bevformer-p150

BEVFormer-small, the camera-only 3D detector of Autoware's autoware_tensorrt_bevformer package (ResNet-101 with deformable convolutions, a 3-layer spatiotemporal BEV encoder over a 150×150 bird's-eye view with the previous frame's BEV as history, a 6-layer DETR decoder), ported to one Tenstorrent Blackhole p150 with tt-nn. Six surround-view camera images with their calibration and the ego pose in; 3D boxes with scores, velocities, the network class and the Autoware label out. Weights: changh95/bevformer-small-weights main (a mirror of the official checkpoint) · Paper: BEVFormer (arXiv:2203.17270) · Autoware package: autoware_tensorrt_bevformer · Training code: fundamentalvision/BEVFormer; export graph: DerryHub/BEVFormer_tensorrt · Port: code/

Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. The whole network runs on the chip as two metal traces replayed back to back (image: the backbone, the neck and the attention value maps; bev: the BEV encoder, the decoder and the heads); the image pre-processing, the camera geometry tables and the box decoding run on the host. All numbers on this card were measured in this configuration.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.

hf download changh95/bevformer-p150 --exclude "image/*" --local-dir bevformer-p150 && cd bevformer-p150
pip install -e .                        # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub, safetensors; ttnn and torch come from tt-metal
pip install -e ".[server,test]"         # optional: the HTTP server and the tests

Run the snippet from the model repo root (the sample path is relative to it).

from tt_bevformer import BEVFormer, load_sample

with BEVFormer.from_pretrained(device_id=0) as model:     # weights -> your HF cache, traces captured
    out = model(**load_sample("code/tt_bevformer/samples/synthetic_6cam.json"))   # six cameras + calibration + stream

for d in out.to_dicts():
    print(d["label"], out.meta["class_names"][d["label_id"]], d["score"], d["center"], d["size"], d["yaw"], d["velocity"])
  • from_pretrained downloads changh95/bevformer-small-weights at the pinned commit 0e6d9d9c7b8 (branch main; the mirror has no tags; 238 MB) to your HF cache, opens the chip, builds the graph and captures the two metal traces. The first load on a machine compiles the kernels (minutes; a cold compile of the WORKER-grid kernels measured 176-178 s); later loads take about 15 s (14.8 s measured).
  • The first call is as fast as the later calls: the load ends with one synthetic frame through the whole path.
  • The shipped sample is a synthetic test pattern with six objects (see code/tt_bevformer/samples/README.md); feed your own rig's six images and calibration for real use.
  • The with block releases the traces and closes the chip. Without with, call model.close().
Input images: the six cameras CAM_FRONT, CAM_FRONT_RIGHT, CAM_FRONT_LEFT, CAM_BACK, CAM_BACK_LEFT, CAM_BACK_RIGHT, any order, any size (path, bytes, PIL, uint8 array; images that are not 1600×900 are resized like the node). calibration: per camera the intrinsics of the image as sent and T_ref_from_camera (camera optical frame -> base_link, or the network's LIDAR_TOP), or a preset. stream: id, reset, timestamp_s, T_world_from_ego (the ego pose that drives the BEV history). ego (optional): CAN dynamics accel, rotation_rate, velocity.
Options score_threshold=0.2 (the node's score_thre), max_num=300 (the node's top-k), physical_boxes=False. from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None).
Output Detections3D in the node's LIDAR_TOP frame: boxes float32 [N, 7] (x, y, z, then the node's size [w, l, h] and yaw), scores, label_ids (the 10 network classes, meta["class_names"]), labels (the Autoware label of each class), velocities [N, 2], timing_ms, meta (the raw network angle and decoder query of every box).
Methods out.to_dict() gives the /predict JSON; out.to_dicts() one dict per box (label, label_id, score, center, size, yaw, velocity).
  • History. Each stream["id"] keeps its previous BEV on the chip; up to 4 at once, a fifth evicts the least recently used one. A new id, reset=True or a timestamp gap above 2 s restarts the history; the history advances once per call, like the node. Pass stream["T_world_from_ego"] (4×4, base_link in a world frame) every frame: the encoder aligns the previous BEV with the ego motion.
  • The API gives the same output as the HTTP server /predict: both share the decoders, the device traces and the host post-processing. One model uses one chip; calls from several threads are serialised.
  • Full reference: code/PYTHON.md. Runnable example: examples/quickstart.py (also writes a bird's-eye view PNG).

Serving (HTTP)

tt-model pull  changh95/bevformer-p150 --with-weights
tt-model serve changh95/bevformer-p150       # or with tt-cli: tt serve changh95/bevformer-p150
S=code/tt_bevformer/samples/synthetic_6cam
python3 code/tt_bevformer/server/client.py --image CAM_FRONT=$S/CAM_FRONT.png --image CAM_FRONT_RIGHT=$S/CAM_FRONT_RIGHT.png \
    --image CAM_FRONT_LEFT=$S/CAM_FRONT_LEFT.png --image CAM_BACK=$S/CAM_BACK.png --image CAM_BACK_LEFT=$S/CAM_BACK_LEFT.png \
    --image CAM_BACK_RIGHT=$S/CAM_BACK_RIGHT.png --calib-preset synthetic_6cam --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/bevformer-p150
  • The image does not contain the weights. --with-weights puts them in your HF cache.
  • The server uses port 20000 (or the next free port). It is ready when the log shows Application startup complete.
  • POST /predict: images (the six cameras, base64 JPEG / PNG / .npy), calibration (per camera intrinsics + T_ref_from_camera, or {"preset": name}); optional stream (id, reset, timestamp_s, T_world_from_ego), ego, params (score_threshold, max_num, physical_boxes), output_format. Also GET /health, GET /info, GET /v1/models (stub). Contract: SERVING.md section 3.

The shipped sample (first request after the warm-up, host server, 2026-10-09):

{"model": "bevformer-p150", "frame_id": "LIDAR_TOP", "num_detections": 6,
 "detections": [
   {"label": "CAR", "label_id": 0, "score": 0.9803, "center": [-0.338, 12.007, -2.927], "size": [1.85, 4.239, 1.512], "yaw": -0.6902, "velocity": [-0.0, 0.003]},
   {"label": "CAR", "label_id": 0, "score": 0.9781, "center": [-3.583, 25.743, -3.899], "size": [2.057, 4.946, 1.894], "yaw": 0.4096, "velocity": [-0.0, 0.002]},
   {"label": "PEDESTRIAN", "label_id": 8, "score": 0.974, "center": [-3.535, 8.988, -2.843], "size": [0.785, 0.846, 1.81], "yaw": 1.4958, "velocity": [-0.008, -0.013]},
   ...],
 "meta": {"class_names": [...], "box_convention": "autoware", "rot": [...], "query": [...], "stream": {...}},
 "timing_ms": {"device_run": 668.5, "preprocess": 358.6, "device": 668.5, "postprocess": 1.3, "total": 1593.7, "decode": 556.1, "model_call": 1036.5}}
  • Boxes are in the network's nuScenes-style LIDAR_TOP frame (x right, y forward, 1.84 m above the road on the default virtual mount), the frame the Autoware node publishes; center is the gravity centre. By default size / yaw follow the node's message: size = [w, l, h] and yaw = -rot + pi (the heading minus 90 degrees); params: {"physical_boxes": true} gives [l, w, h] and the heading of the length axis, and meta.rot the raw network angle.
  • label is the Autoware label the node publishes (bus, trailer and construction_vehicle become TRUCK; barrier and traffic_cone UNKNOWN); label_id is the network class (meta.class_names).
  • velocity is [vx, vy] in m/s in the LIDAR_TOP frame, as trained.
  • /info reports the device (ETH, 12×10), the labels, the stream table (4 streams) and the load-time knobs.

Demo

PandaSet 019, frame 40 (San Francisco; CC BY 4.0): TT output nuScenes scene-0103, key-frame 20 (Boston; CC BY-NC-SA 4.0, non-commercial): TT output
PandaSet 019 frame 40, six cameras and bird's-eye view nuScenes scene-0103 key-frame 20, non-commercial
TT vs the fp32 CPU reference: PandaSet 019 at 2 Hz The shipped synthetic sample: TT vs the stored CPU reference
TT vs CPU on PandaSet 019 The synthetic sample

More: nuScenes scene-0916 key-frame 20 (Singapore) and the key-frames of scene-0103 as a GIF, both non-commercial; PandaSet 019 at 2 Hz as a GIF. Every image shows outputs of this port on the p150 (the boxes the node publishes, score > 0.2), some next to the fp32 CPU reference; grey dashed outlines are the dataset annotations. Sources, licences and changes: media/ATTRIBUTION.md. The weights were trained on nuScenes; PandaSet is another rig (a narrower front camera, wider side cameras, mounted higher), so its detections show the domain gap, not the port: on frame 40 only 22 of the 63 published boxes match one of the 78 annotations within 50 m (one-to-one, same class, < 2 m).

Demo & Performances

Warm, batch 1, six cameras. Accuracy: the TT output against the port's fp32 CPU reference of the same network (same weights, same pre- and post-processing), which matches ONNX Runtime on a plugin-free ONNX of the export graph (PCC 1.0000000000). Speed: code/scripts/bench.py, p50 (p99) of 60 iterations, on a shared 8-core host (load average 4.3-5.6; OPT_BASELINE.md).

Metric Performance
Per-stage PCC vs the CPU reference, PandaSet frame 35 (gates) backbone 0.99941, neck 0.99979 (gate 0.999); encoder layers 0.99999 / 0.99998, decoder 0.99997 (0.999, teacher-forced); bev_embed 0.99986 (0.999); last-layer classes 0.99978 (0.995), boxes 0.999994 (0.9999)
nuScenes mini_val mAP / NDS (81 key-frames; indicative: 3 of 10 classes absent) test-time decode (top-300 boxes): TT 0.3539 / 0.3985, CPU 0.3527 / 0.3984 (the checkpoint's full-val result: 0.370 / 0.479); the node's published boxes (score > 0.2): TT 0.3398 / 0.3949, CPU 0.3392 / 0.3950
Box agreement, nuScenes mini_val (all 81 key-frames, published boxes) recall 0.986, precision 0.995 vs the CPU run (same class, 1 m)
Box agreement, PandaSet 019 frames 30 -> 35 -> 40 through model() recall 0.966, precision 1.000, max |score difference| 0.034 (gates 0.95 / 0.95 / 0.05)
Box agreement, the shipped synthetic sample (the container smoke's check) 6 / 6 boxes (recall 1.000, precision 1.000), max |score difference| 0.0005 vs the stored CPU reference (gates 0.95 / 0.95 / 0.05)
Device, one frame (image + bev replays back to back) 485.6 ms: 2.06 frames/s device-bound
Trace replays: image / bev 243.8 / 240.8 ms (synthetic input), 245.2 / 240.7 ms (PandaSet)
Python model(), 1600×900 images (synthetic) 939 ms (965): preprocess 303, image upload 141, device 486, postprocess 1.4
Python model(), PandaSet 1920×1080 (host resize to 1600×900) 1,555 ms (1,799): the host's uint8 resize alone 603 ms
Served /predict on the host, shipped sample (uvicorn, loopback) one request after the warm-up: timing_ms.total 1,594 ms (decode of the 13.7 MB PNG body 556, preprocess 359, device incl. upload 669, postprocess 1.3); client round trip 1,860 ms

All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, AICLK 1350 MHz. The device is kernel-bound (485.1 ms of kernels; 0.74 ms of op-to-op gaps for 1,317 programs), and half of the device time is grid_sample: the deformable-attention sampling of the encoder (178 ms) and the deformable convolutions of the backbone (77 ms). The end-to-end latency adds the host pre-processing and the 34 MB image upload. This is the first release: optimization has not started (OPT_REPORT.md ranks the plan; OPT_BASELINE.md has every measurement). Verification: VERIFICATION_2026-10-09.md.

The nuScenes metrics use the official devkit (detection_cvpr_2019, mini_val: scene-0103 and scene-0916, history reset at each scene start, 2 Hz key-frames), once on BEVFormer's test-time decode (the top 300 boxes, no score threshold: the protocol of the full-val result) and once on the boxes the node publishes (score > 0.2). nuScenes dataset © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use); non-commercial: the data were used for local validation only and are not distributed here; Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020. Box agreement: a reference box counts as recalled when a TT box of the same class lies within the stated distance, and the reverse for precision. On PandaSet the TT scores sit slightly below the CPU's (mean -0.002 to -0.003), so a box within a few hundredths of the 0.2 threshold can drop out (2 of 58 on frame 35).

No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). The Autoware pull request that added the package (autoware_universe #11076) reports a 91 ms TensorRT FP16 engine time on an RTX 2080 Ti and about 303 ms end to end including the host pre-processing (not like-for-like). p150 power was not measured, so no efficiency comparison is made.

Caveats

  • Deployment status. autoware_tensorrt_bevformer is a standalone package: no Autoware perception mode or launch file starts it (the default perception is LiDAR CenterPoint). This port is a drop-in for the network and the node's pre- and post-processing, not a ROS node, and not a certified Autoware component: do not use it for safety-critical driving decisions.
  • Autoware's ONNX is not available. The node loads a bevformer_small.onnx whose download link is archived. This port runs the official bevformer_small_epoch_24 checkpoint (mirrored, byte-identical tensors) through a re-implementation of the graph of DerryHub's TensorRT export, the export the Autoware model comes from. Whether Autoware's file is the nv_half or the nv_half2 variant of that export is not known; both have the same graph.
  • First release: optimization pending. This is the functional baseline port; no optimization round has run yet (OPT_REPORT.md).
  • Precision. The node defaults to an fp16 TensorRT engine; this port follows the fp32 graph. On the chip: fp32 weights and biases, HiFi4, fp32 and packer L1 accumulation, bf16 activations, and the deformable-attention sampling with fp32 grids and weights (the precision policy of the first release). Outputs differ slightly from the fp32 reference (agreement figures above).
  • Documented deviations from the node (all of them selectable or explained in SERVING.md):
    • K is rescaled to the node's 1600×900 (the node keeps the raw CameraInfo.K when it resizes; calibration.scale_intrinsics: false reproduces it).
    • The ego dynamics fill can_bus in BEVFormer's training layout; the node's literal KinematicState order (BEVFORMER_CAN_BUS_LAYOUT=autoware) swaps velocity and acceleration. Without an ego field both give the same input.
    • The history restarts on a new stream id, reset or a gap above 2 s (the node never resets); in one steady 10 Hz stream that never happens, so the outputs are the node's.
    • A rig whose cameras each see more than 6,144 BEV cells is refused with a 400 (the static capacity of the spatial cross-attention rebatching); the node's dense graph accepts any rig. Every rig tested fits (at most 5,704 cells).
  • Temporal cadence. The network has no Δt input: it was trained on key-frames 0.5-1.0 s apart, and the node advances the history at camera rate. Send frames at about 2 Hz for the training regime.
  • Domain gap. The weights were trained on nuScenes (Boston, Singapore). The network assumes a nuScenes-like rig and a LIDAR_TOP frame about 1.84 m above the road; other rigs need a virtual LIDAR_TOP (the default T_ref_from_lidar) and lose accuracy (PandaSet 019 frame 40: of the 63 published boxes, 22 match one of the 78 annotations within 50 m, one-to-one, same class, < 2 m).
  • Fixed shapes. Six cameras in the nuScenes slots, 1280×736 network input (1600×900 ×0.8), a 150×150 BEV grid (±51.2 m), 900 queries, 10 classes; batch 1, one frame per request; requests are serialised on the chip; up to 4 stream histories.
  • Validation scope. Agreement with the fp32 CPU reference on public data (nuScenes mini_val, PandaSet) and on the synthetic sample; the mini_val mAP / NDS are indicative, not a full-val reproduction. The shipped sample is synthetic; the PandaSet sample is not shipped yet.
  • Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication.
  • dispatch="worker" (server: BEVFORMER_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (503.5 ms per frame); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.

Licensing

  • Weights: changh95/bevformer-small-weights at branch main (commit 0e6d9d9c7b8aca95820b264491e19a72d10f0608), Apache-2.0 per its model card: the state_dict of the official bevformer_small_epoch_24.pth as safetensors, with attribution to the BEVFormer authors, its source URL and sha256 (the authors publish it as a release asset linked from the Apache-2.0 fundamentalvision/BEVFormer README; the asset itself carries no licence file). Not redistributed here: the package only points to them. The mirror card's training-data notice: the weights were trained on nuScenes (CC BY-NC-SA 4.0, non-commercial; commercial use of nuScenes needs a licence from Motional); read it before commercial use.
  • Pre- and post-processing ported from autoware_universe perception/autoware_tensorrt_bevformer (Apache-2.0). The network graph follows DerryHub/BEVFormer_tensorrt (Apache-2.0); no code was copied from it.
  • Port and serving code (code/): Apache-2.0.
  • Sample data: code/tt_bevformer/samples/synthetic_6cam* and the preset calib/synthetic_6cam.json were generated by this repository (code/scripts/make_sample.py; no third-party data), Apache-2.0. No dataset sample is shipped; nuScenes, Argoverse 2 and KITTI never ship, and a PandaSet sample (CC BY 4.0) is prepared but not yet included.
  • Demo media (media/; per-file sources, frames and changes: media/ATTRIBUTION.md):
    • PandaSet renders (media/bevformer_pandaset_*): Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: camera images downscaled; predicted 3D boxes, annotation outlines and LiDAR points drawn on top; BEV plots. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021.
    • nuScenes renders (media/*_NC.*), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020.
    • The synthetic-sample render (media/bevformer_synthetic_6cam.png): generated by this repository, Apache-2.0.

Provenance

These are the exact sources the container image was built from:

component built from
tt-metal 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (dirty tree: the image includes the patch)
weights changh95/bevformer-small-weights@0e6d9d9c7b8aca95820b264491e19a72d10f0608 (branch main; the mirror has no tags), files bevformer_small_epoch_24.safetensors (sha256 51ba3128…95d7; 1,001 tensors bit-identical to the official bevformer_small_epoch_24.pth, sha256 4ebc5810…4bc5), provenance.json
Autoware reference autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (package 0.53.0); the graph of DerryHub's forward_trt export, re-implemented
shared port code ttaw 0.20.0 @ common 89dec49 (vendored in code/tt_bevformer/ttaw, VENDORED.json)
code/ digest (image) 63465fcf45ce26ac (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json)
image tt-model/bevformer-p150:30a017307fbe (sha256:30a017307fbe0e054512d1090af266bdc04dac45c525957b85cb5b6c7ebb84a3)
base images build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json)
built 2026-10-09T12:07:02+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/bevformer-p150

Finetuned
(1)
this model

Collection including changh95/bevformer-p150

Paper for changh95/bevformer-p150