sceneseg-p150
SceneSeg (Autoware VisionPilot): the scene segmentation network of the Autoware Foundation's VisionPilot camera stack, as its scene_seg_model node deployed it, ported to one Tenstorrent Blackhole p150 with tt-nn. One camera image of any size in (converted to BGR8, as the node receives it); the VisionPilot foreground mask (mono8, 255 = foreground, at the input's size) and a 3-class map (background / foreground / road) out, with the node's exact pre- and post-processing.
Weights: weights/ of this repo @ a6829b24 (VisionPilot's SceneSeg_FP32.onnx and its .pth as safetensors, re-hosted from the upstream Google Drive links) · Paper: none (model card AutowareFoundation/SceneSeg) · Autoware package: VisionPilot models (node scene_seg_model) · Training code: autowarefoundation/vision_pilot Models/ @ ca58cb50 · Port: code/
Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Two serve profiles, the input channel order the network is fed: bgr (the deployed VisionPilot order, the default) and rgb (the training order). Numerics: weights as two bf16 terms (hi + lo), backbone activations in float32 fed as two bf16 terms, biases added in fp32, fp32 accumulation. All numbers on this card were measured in this configuration (bgr unless stated).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/sceneseg-p150 --exclude "image/*" --exclude "weights/*" --local-dir sceneseg-p150 && cd sceneseg-p150
pip install -e . # adds numpy<2, pillow, pyyaml, onnx, safetensors, huggingface_hub; ttnn and torch come from tt-metal
pip install -e ".[server,test]" # optional: the HTTP server and the tests
Run the snippet from the model repo root: code/tt_sceneseg/samples/highway_normal_1.png is a path relative to it.
from tt_sceneseg import SceneSeg
with SceneSeg.from_pretrained(device_id=0) as model: # weights -> your HF cache, trace captured
out = model("code/tt_sceneseg/samples/highway_normal_1.png") # path, bytes, PIL image or RGB uint8 HxWx3 array of any size
mask = out.mask # uint8 (H, W) at the input size: 255 = foreground (the VisionPilot topic)
print(out.class_counts()) # pixels per class of out.class_map (background / foreground / road)
from_pretraineddownloadsweights/SceneSeg_FP32.onnx(+ LICENSE, NOTICE, provenance; 194 MB) of this repo at the pinned commita6829b246a0to your HF cache, opens the chip, builds the graph and captures the metal trace. With a warm kernel cache the load takes about 13 s; the first load on a machine also compiles the kernels (minutes: a configuration whose kernels were not cached yet took 166-223 s of warm-up here).- The trace is captured during the load, so the first call is as fast as the later ones.
- The
withblock releases the trace and closes the chip. Withoutwith, callmodel.close().
| Input | One camera image: path, PNG / JPEG bytes, PIL image, RGB uint8 HxWx3 array (any size), or images=[...] like the server. Converted to BGR8 and squashed to 640x320 with OpenCV's exact INTER_LINEAR (no letterbox, aspect ratio not kept), as the node does. An OpenCV / ROS bgr8 array must be flipped first (bgr[:, :, ::-1]). |
| Options | class_map=True. from_pretrained(device_id=0, variant="bgr" | "rgb", dispatch="eth", weights_dir=None, device=None): variant is the channel order the network is fed (see "Demo & Performances", channel order). |
| Output | SceneSegMask (a ttaw.outputs.Mask2D): mask uint8 (H, W) 0 / 255 at the input size, class_map uint8 (H, W) 0 / 1 / 2, class_names, meta, timing_ms. |
| Methods | out.to_dict() gives the /predict JSON; out.class_counts() the pixels per class of the class map. |
- The API gives the same output as the HTTP server
/predict: both share the decoders, the device trace and the host post-processing (checked on the device bytest_api_equals_server). - One model uses one chip; calls from several threads are serialised.
- Full reference:
PYTHON.md. Runnable example:examples/quickstart.py(also writesquickstart_mask.pngandquickstart.png, the class map over the image).
Serving (HTTP)
tt-model pull changh95/sceneseg-p150 --with-weights
tt-model serve changh95/sceneseg-p150 # or with tt-cli: tt serve changh95/sceneseg-p150; --profile rgb for the training order
python3 code/tt_sceneseg/server/client.py --image CAM_FRONT=code/tt_sceneseg/samples/highway_normal_1.png --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/sceneseg-p150
- The image does not contain the weights.
--with-weightsputs them in your HF cache. - The server uses port 20000 (or the next free port). It is ready when the log shows
Application startup complete(15 s with a warm kernel cache, measured with uvicorn on the host). - Two serve profiles (
tt-model profiles changh95/sceneseg-p150):bgr(default) andrgb. client.pybuilds the request with the standard library only; add--url http://127.0.0.1:20000to send it.POST /predict:images(one entry: base64 PNG / JPEG of the camera image, any size, decoded to BGR8 like the node's input), optionalparams(class_maptrue),output_format. AlsoGET /health,GET /info,GET /v1/models(stub). Contract:SERVING.mdsection 3.
The served body for the shipped sample (bgr profile; PNG data elided):
{"model": "sceneseg-p150", "frame_id": "camera",
"meta": {"channel_order": "bgr", "source_hw": [414, 727], "network_hw": [320, 640], "foreground_class": 1, "mask_value": 255, "foreground_fraction": 0.040146},
"mask": {"format": "png", "key": "mask", "dtype": "uint8", "shape": [414, 727], "data": "<base64 PNG>"},
"class_names": ["background", "foreground", "road"],
"class_counts": {"background": 166572, "foreground": 12083, "road": 122323},
"class_map": {"format": "png", "key": "class_map", "dtype": "uint8", "shape": [414, 727], "data": "<base64 PNG>"},
"timing_ms": {"preprocess": 14.66, "device": 67.03, "postprocess": 1.27, "total": 113.84, "decode": 16.64, "model_call": 82.98}}
maskis what the VisionPilot node publishes on/autoseg/scene_seg/mask: mono8 at the input's size, 255 where the per-pixel argmax is class 1 (foreground), else 0, resized back from 640x320 with nearest neighbour.class_mapis the 3-class argmax on the same grid (0 background incl. sky, 1 foreground objects, 2 drivable road), a by-product the node does not publish;class_countscounts its pixels.
Demo
Input (code/tt_sceneseg/samples/highway_normal_1.png, the shipped sample) |
Foreground mask and road on p150 (media/sceneseg_highway_normal_1_tt.jpg) |
|---|---|
![]() |
![]() |
The p150 output next to the fp32 CPU reference on the same image; the bottom panel marks in white the pixels where the two class maps differ (24 of 300,978):
All nine shipped VisionPilot tutorial images, and three of the shipped comma10k images against their ground truth:
On public driving datasets (p150 outputs; the frames themselves are not in this repository). The caption of each image gives its agreement with the fp32 CPU reference (KITTI: the foreground IoU against the ground truth).
PandaSet renders: contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms; resized and annotated with model outputs; Scale AI and Hesai do not endorse this work. The nuScenes render is non-commercial (CC BY-NC-SA 4.0): rendered from the nuScenes dataset, © Motional AD Inc., nuScenes Terms of Use; Motional does not endorse this work. The KITTI render is non-commercial (CC BY-NC-SA 3.0): images and labels from the KITTI Vision Benchmark Suite. Sources, changes and the full attributions: media/ATTRIBUTION.md.
Demo & Performances
Warm, batch 1, 2026-10-09. Latency: the stage bench of OPT_BASELINE.md (code/scripts/bench.py, 100 iterations per stage) on the shipped sample code/tt_sceneseg/samples/highway_normal_1.png (727x414) and on a PandaSet front-camera frame (1920x1080, not shipped), re-checked on the release code; the served rows from uvicorn on the host (the app the container runs) and a loopback client, 50 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load; the device rows repeat to within 0.03 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network (same weights, same pre- and post-processing) on the shipped samples, on 23 public gate frames and on 415 held-out public frames.
| Metric | Performance |
|---|---|
| Agreement with the fp32 CPU reference, shipped sample (320x640 class map / published mask) | 99.993 % of pixels / 99.994 %; logits PCC 0.9999982 |
| Agreement, 23 public gate frames (9 tutorial images, 2 comma10k, 8 PandaSet, 4 nuScenes) | logits PCC ≥ 0.9999952; argmax agreement ≥ 99.943 % (mean 99.983 %) |
| Agreement, 415 held-out public frames (KITTI 200, comma10k val 100, nuScenes 55, PandaSet 48, comma10k 12) vs ONNX Runtime | argmax agreement ≥ 99.752 % (mean 99.980 %, 1st percentile 99.852 %); rgb ≥ 99.812 % |
| Ground-truth foreground IoU, the 6 shipped comma10k images (training ROI; training data) | 0.7288 on p150 vs 0.7287 fp32 CPU (rgb 0.7645 vs 0.7647) |
| Module PCC vs the fp32 reference (backbone stages, SceneContext, neck, head, logits; each module fed the reference inputs) | ≥ 0.99995 (gate 0.999) |
Python model() call, shipped sample 727x414 (decode + squash, H2D, trace, D2H, post) |
90.85 ms p50 (p99 91.55) · 11.0 frames/s |
Python model() call, PandaSet 1920x1080 |
116.12 ms p50 (p99 119.01) · 8.6 frames/s |
Served /predict timing_ms.total (uvicorn on the host, shipped sample; includes the base64 + PNG decode, 16.6 ms) |
113.4 ms median (min 112.6) |
| Served client round trip, loopback (467 kB base64 request) | 120.2 ms median |
| Device trace, one blocking forward | 61.14 ms |
| Back-to-back trace replays | 61.06 ms per frame · 16.4 frames/s |
| Host decode · squash · H2D · D2H · host post (shipped sample) | 10.1 · 12.2 · 5.3 · 0.5 · 1.4 ms |
from_pretrained load, warm kernel cache |
13.1 s |
The rgb profile's stage-bench p50s are within 0.35 ms of bgr on every row (the device rows within 0.02 ms); served, its timing_ms.total median was 114.6 ms (20 requests). All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, with the published numerics (two-term bf16 weights, float32 backbone activations as two bf16 terms). Accuracy is agreement with the fp32 CPU reference of the same network; SceneSeg has no paper and no published benchmark, so the dataset metrics below are this port's own measurements. Details: VERIFICATION_2026-10-09.md, OPT_BASELINE.md, OPT_REPORT.md.
Channel order. VisionPilot's deployed C++ pre-processing feeds B, G, R planes (with the ImageNet constants reordered to match) to a network trained on R, G, B: a channel-order bug in the deployment. The default bgr profile reproduces the deployed node; rgb feeds the training order. Both run the same device graph (the order is folded into the first conv's weights) and both are validated against the fp32 CPU reference of their own order. The training order is clearly better, most of all on two-wheelers (foreground IoU unless stated; the datasets are held out except comma10k):
| data (ground truth) | bgr p150 (fp32 CPU) |
rgb p150 (fp32 CPU) |
|---|---|---|
| comma10k validation split, the authors' own protocol (100 images, training ROI; per-image smoothed IoU) | 0.553 (0.553) | 0.601 (0.601) |
| KITTI semantics 2015 (200 frames, 2:1 centre crop; not in the training data; pooled) | 0.785 (0.785) | 0.814 (0.813) |
| PandaSet front camera, LiDAR labels projected (48 frames, 434,527 points) | 0.821 (0.822) | 0.881 (0.881) |
PandaSet: share of Bicycle points predicted foreground (1,886 points) |
0.056 (0.056) | 0.790 (0.791) |
| nuScenes CAM_FRONT, LiDAR labels projected (55 key-frames) | 0.697 (0.697) | 0.713 (0.713) |
nuScenes: share of motorcycle points predicted foreground (59 points) |
0.29 (0.29) | 0.71 (0.73) |
| the 6 shipped comma10k images, training ROI (5 are training images) | 0.729 (0.729) | 0.764 (0.765) |
The p150 column scores the served trace's class maps of all 952 public-frame runs with the same research protocols as
the CPU (research/sceneseg/public_data, ss_sanity.py --cls-dir); the dataset labels are mapped to SceneSeg's
three classes (vehicles, two-wheelers and people are foreground). The LiDAR rows are a sparse proxy (points 1-40 m
projected into the camera), and the comma10k validation split is reconstructed from the upstream code.
Use bgr to reproduce the Autoware VisionPilot output, rgb for better masks.
No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). VisionPilot ran this network as a TensorRT FP16 engine built from SceneSeg_FP32.onnx (autoseg.yaml: backend tensorrt, precision fp16); the upstream model README (Models/model_library/SceneSeg/README.md @ ca58cb50) reports 18.1 FPS FP32 / 26.7 FPS FP16 on an RTX 3060 Mobile (not like-for-like). p150 power was not measured, so no efficiency comparison is made.
Caveats
- First release: baseline port, optimization pending. The device graph is one metal trace and is kernel-bound, but only 41 % of its 61 ms is conv / matmul compute: the rest is the fp32 glue of the precision policy and data movement.
OPT_REPORT.mdranks what comes next. - Deployment status in Autoware: SceneSeg is not part of Autoware Universe. It was deployed by the Autoware Foundation's VisionPilot stack (formerly autoware.privately-owned-vehicles) through its generic ROS 2
modelspackage, nodescene_seg_model(TensorRT FP16; ONNX Runtime and Zenoh wrappers shared the same pre- and post-processing). Upstream removed that deployment code on 2026-06-15 (last tree04fa3e80, which this port follows) and the PyTorch model code on 2026-07-06 (last treeca58cb50); the current VisionPilot 1.0 ships AutoSpeed, AutoSteer and AutoDrive, not SceneSeg. This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions. - Channel order: the default
bgrreproduces the deployed node, including its channel-order bug (above).rgbis the training order, offered as a load-time option; it is not what VisionPilot ran. - Numerics: validated against the fp32 CPU reference of the same graph (ONNX Runtime on the deployed ONNX), not against VisionPilot's TensorRT FP16 engine.
- Precision policy of this release: every conv, transposed conv and linear runs its weights as two bf16 terms (hi + lo) in separate bias-free convs summed in fp32; the backbone keeps float32 activations and feeds each layer two bf16 terms of them; every bias is added as an fp32 row (a bias fused into a ttnn conv or linear truncates its fp32 result toward zero); HiFi4, fp32 accumulation, packer L1 accumulation; SceneContext and the decoder run bf16 activations. Plain bf16 weights replay in 39.5 ms instead of 61.1 ms, but they missed the 99.5 % agreement gate on a nuScenes frame (0.9949, before fix round 1), and bf16 backbone activations miss it on 3 held-out frames in the CPU emulation (
OPT_REPORT.md, "Rejected"). A small residual remains: the stock depthwise and 3x3 convs keep a slight toward-zero bias even without a fused bias, so the hardest held-out frame stays at 99.75 % agreement. - Documented deviations from the node: the squash to 640x320 (bit-exact with
cv::resizeINTER_LINEAR) and the INTER_NEAREST resize back run on the host, the normalisation, the network and the class decision (argmax, lowest index on ties, as the node's strict>scan) on the device; the 3-class map is returned as well; the mask is a PNG (or npz) in JSON, not a ROS image message; inputs are RGB arrays or encoded files (converted to BGR8 inside). SceneSeg's Python tooling also offered a bicubic resize; it is not offered here, because the deployed node never used it. The INT8 QAT ONNX (never deployed) and SceneSegLite (a different 19-class network) are not part of this package. - Domain: SceneSeg was trained on ACDC, MUSES, IDDAW, Mapillary Vistas and comma10K (roughly 2:1 views without the ego hood) for a front camera with 52-55° HFOV. On the public frames used here the fp32 CPU reference itself shows the weights' limits, and the p150 output agrees with it: small and distant people and two-wheelers are the weak spot (above all in the deployed
bgrorder), parking areas and a third of sidewalk points come out as road, trams stay background, and wider cameras (nuScenes 65°) or frames with the ego hood (raw comma10k) score lower. - Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (
patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication. dispatch="worker"(server:SCENESEG_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (63.82 ms per replay instead of 61.06 ms); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.- Batch 1, one frame per request; requests are serialised on the chip. The network input is fixed at 640x320 (the SceneContext reshape hard-codes the 10x20 stride-32 grid); any source image size is accepted, squashed on the host without keeping the aspect ratio, and the mask is resized back with nearest neighbour.
- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150 power was not measured, so no efficiency comparison is made.
Licensing
- Weights:
weights/of this repo at commita6829b246a05dba55c356323bbf75f7b1d148499: VisionPilot'sSceneSeg_FP32.onnx(Google Drive1l-dniunvYyFKvLD7k16Png3AsVTuMl9f, sha2568e509094…c973) and its.pth(Drive1vCZMdtd8ZbSyHn1LCZrbNKMK7PQvJHxj, sha2569482f131…89fc) converted to safetensors, re-hosted unchanged with the upstream Apache-2.0 LICENSE, NOTICE and provenance.json (the upstream README covers the model weights with Apache-2.0; the card AutowareFoundation/SceneSeg holds no weights). The upstream model card lists the training data: ACDC, MUSES, IDDAW, Mapillary Vistas and comma10K (BDD100K held out); check those datasets' terms before commercial use. - Pre- and post-processing ported from autowarefoundation/vision_pilot
VisionPilot/middleware_recipes/ROS2/modelsandmiddleware_recipes/common/backends@04fa3e80(Apache-2.0; removed upstream on 2026-06-15). - Port and serving code (
code/): Apache-2.0.patches/tt-metal-eth-dispatch.patchmodifies tt-metal (Apache-2.0). - Sample data (only redistributable data ships; the public-dataset frames of the accuracy tables are not in this repository):
code/tt_sceneseg/samples/*.png: the 9 images of the upstream SceneSeg tutorial (autowarefoundation/vision_pilot @ca58cb50,Models/tutorials/assets/images/), Apache-2.0 (samples/LICENSE-vision_pilot.txt); the original photo source is not stated upstream.code/tt_sceneseg/samples/comma10k/: six comma10k images with their masks (commaai/comma10k @6c205fe), MIT, © 2020 Comma.ai, Inc. (samples/comma10k/LICENSE-comma10k.txt); five of them are SceneSeg training images.
- Demo media (
media/, sources and changes inmedia/ATTRIBUTION.md):- PandaSet renders: Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: resized, annotated with model outputs. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021.
- nuScenes render (
media/sceneseg_nuscenes_scene0103_k28_front_tt_NC.jpg), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020. - KITTI render (
media/sceneseg_kitti_heldout_tt_NC.jpg), non-commercial, CC BY-NC-SA 3.0: Contains images and labels from the KITTI Vision Benchmark Suite, semantic segmentation benchmark 2015 (A. Geiger, P. Lenz, C. Stiller, R. Urtasun; H. Alhaija et al.), https://www.cvlibs.net/datasets/kitti/. Non-commercial use only; adaptations under the same license. Changes: cropped to 2:1, resized, prediction and ground truth overlaid. Cite: A. Geiger et al., CVPR 2012; H. Alhaija et al., IJCV 2018; M. Menze and A. Geiger, CVPR 2015. - Tutorial-image renders: Apache-2.0; comma10k renders: MIT.
Provenance
These are the exact sources the container image was built from:
| component | built from |
|---|---|
| tt-metal | 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (sha256 08d0ddf6…; dirty tree: the image includes the patch) |
| weights | changh95/sceneseg-p150@a6829b246a05dba55c356323bbf75f7b1d148499, files weights/SceneSeg_FP32.onnx, weights/provenance.json, weights/LICENSE, weights/NOTICE (ONNX sha256 8e509094…c973, checked at load) |
| Autoware reference | autowarefoundation/vision_pilot 04fa3e80d0b0a9d2491c07991a9cb48101b2c2cb (deployment, VisionPilot/middleware_recipes); model code ca58cb50 |
| shared package | ttaw 0.22.0, vendored as code/tt_sceneseg/ttaw from the Autoware ports' shared common repository at commit 2218be3 (code/tt_sceneseg/ttaw/VENDORED.json: version, commit and per-file sha256) |
code/ digest (image) |
7bbcad160628eb66 (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json) |
| image | tt-model/sceneseg-p150:f070810e48b2 (sha256:f070810e48b2785605cef8dc572be0ddfa4db141ce47c93c71266a0b581f9a18) |
| base images | build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json) |
| built | 2026-10-09T18:40:28+00:00 by tt-model 0.1.0 |











