sceneseg-p150

SceneSeg (Autoware VisionPilot): the scene segmentation network of the Autoware Foundation's VisionPilot camera stack, as its scene_seg_model node deployed it, ported to one Tenstorrent Blackhole p150 with tt-nn. One camera image of any size in (converted to BGR8, as the node receives it); the VisionPilot foreground mask (mono8, 255 = foreground, at the input's size) and a 3-class map (background / foreground / road) out, with the node's exact pre- and post-processing. Weights: weights/ of this repo @ a6829b24 (VisionPilot's SceneSeg_FP32.onnx and its .pth as safetensors, re-hosted from the upstream Google Drive links) · Paper: none (model card AutowareFoundation/SceneSeg) · Autoware package: VisionPilot models (node scene_seg_model) · Training code: autowarefoundation/vision_pilot Models/ @ ca58cb50 · Port: code/

Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Two serve profiles, the input channel order the network is fed: bgr (the deployed VisionPilot order, the default) and rgb (the training order). Numerics: weights as two bf16 terms (hi + lo), backbone activations in float32 fed as two bf16 terms, biases added in fp32, fp32 accumulation. All numbers on this card were measured in this configuration (bgr unless stated).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.

hf download changh95/sceneseg-p150 --exclude "image/*" --exclude "weights/*" --local-dir sceneseg-p150 && cd sceneseg-p150
pip install -e .                        # adds numpy<2, pillow, pyyaml, onnx, safetensors, huggingface_hub; ttnn and torch come from tt-metal
pip install -e ".[server,test]"         # optional: the HTTP server and the tests

Run the snippet from the model repo root: code/tt_sceneseg/samples/highway_normal_1.png is a path relative to it.

from tt_sceneseg import SceneSeg

with SceneSeg.from_pretrained(device_id=0) as model:      # weights -> your HF cache, trace captured
    out = model("code/tt_sceneseg/samples/highway_normal_1.png")   # path, bytes, PIL image or RGB uint8 HxWx3 array of any size

mask = out.mask                    # uint8 (H, W) at the input size: 255 = foreground (the VisionPilot topic)
print(out.class_counts())          # pixels per class of out.class_map (background / foreground / road)
  • from_pretrained downloads weights/SceneSeg_FP32.onnx (+ LICENSE, NOTICE, provenance; 194 MB) of this repo at the pinned commit a6829b246a0 to your HF cache, opens the chip, builds the graph and captures the metal trace. With a warm kernel cache the load takes about 13 s; the first load on a machine also compiles the kernels (minutes: a configuration whose kernels were not cached yet took 166-223 s of warm-up here).
  • The trace is captured during the load, so the first call is as fast as the later ones.
  • The with block releases the trace and closes the chip. Without with, call model.close().
Input One camera image: path, PNG / JPEG bytes, PIL image, RGB uint8 HxWx3 array (any size), or images=[...] like the server. Converted to BGR8 and squashed to 640x320 with OpenCV's exact INTER_LINEAR (no letterbox, aspect ratio not kept), as the node does. An OpenCV / ROS bgr8 array must be flipped first (bgr[:, :, ::-1]).
Options class_map=True. from_pretrained(device_id=0, variant="bgr" | "rgb", dispatch="eth", weights_dir=None, device=None): variant is the channel order the network is fed (see "Demo & Performances", channel order).
Output SceneSegMask (a ttaw.outputs.Mask2D): mask uint8 (H, W) 0 / 255 at the input size, class_map uint8 (H, W) 0 / 1 / 2, class_names, meta, timing_ms.
Methods out.to_dict() gives the /predict JSON; out.class_counts() the pixels per class of the class map.
  • The API gives the same output as the HTTP server /predict: both share the decoders, the device trace and the host post-processing (checked on the device by test_api_equals_server).
  • One model uses one chip; calls from several threads are serialised.
  • Full reference: PYTHON.md. Runnable example: examples/quickstart.py (also writes quickstart_mask.png and quickstart.png, the class map over the image).

Serving (HTTP)

tt-model pull  changh95/sceneseg-p150 --with-weights
tt-model serve changh95/sceneseg-p150       # or with tt-cli: tt serve changh95/sceneseg-p150; --profile rgb for the training order
python3 code/tt_sceneseg/server/client.py --image CAM_FRONT=code/tt_sceneseg/samples/highway_normal_1.png --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/sceneseg-p150
  • The image does not contain the weights. --with-weights puts them in your HF cache.
  • The server uses port 20000 (or the next free port). It is ready when the log shows Application startup complete (15 s with a warm kernel cache, measured with uvicorn on the host).
  • Two serve profiles (tt-model profiles changh95/sceneseg-p150): bgr (default) and rgb.
  • client.py builds the request with the standard library only; add --url http://127.0.0.1:20000 to send it.
  • POST /predict: images (one entry: base64 PNG / JPEG of the camera image, any size, decoded to BGR8 like the node's input), optional params (class_map true), output_format. Also GET /health, GET /info, GET /v1/models (stub). Contract: SERVING.md section 3.

The served body for the shipped sample (bgr profile; PNG data elided):

{"model": "sceneseg-p150", "frame_id": "camera",
 "meta": {"channel_order": "bgr", "source_hw": [414, 727], "network_hw": [320, 640], "foreground_class": 1, "mask_value": 255, "foreground_fraction": 0.040146},
 "mask": {"format": "png", "key": "mask", "dtype": "uint8", "shape": [414, 727], "data": "<base64 PNG>"},
 "class_names": ["background", "foreground", "road"],
 "class_counts": {"background": 166572, "foreground": 12083, "road": 122323},
 "class_map": {"format": "png", "key": "class_map", "dtype": "uint8", "shape": [414, 727], "data": "<base64 PNG>"},
 "timing_ms": {"preprocess": 14.66, "device": 67.03, "postprocess": 1.27, "total": 113.84, "decode": 16.64, "model_call": 82.98}}
  • mask is what the VisionPilot node publishes on /autoseg/scene_seg/mask: mono8 at the input's size, 255 where the per-pixel argmax is class 1 (foreground), else 0, resized back from 640x320 with nearest neighbour.
  • class_map is the 3-class argmax on the same grid (0 background incl. sky, 1 foreground objects, 2 drivable road), a by-product the node does not publish; class_counts counts its pixels.

Demo

Input (code/tt_sceneseg/samples/highway_normal_1.png, the shipped sample) Foreground mask and road on p150 (media/sceneseg_highway_normal_1_tt.jpg)

The p150 output next to the fp32 CPU reference on the same image; the bottom panel marks in white the pixels where the two class maps differ (24 of 300,978):

All nine shipped VisionPilot tutorial images, and three of the shipped comma10k images against their ground truth:

The 9 tutorial images (p150) comma10k: p150 vs ground truth (training data of SceneSeg)

On public driving datasets (p150 outputs; the frames themselves are not in this repository). The caption of each image gives its agreement with the fp32 CPU reference (KITTI: the foreground IoU against the ground truth).

PandaSet 019, frame 40, front camera PandaSet 019, 8 s at 10 Hz (every 2nd frame)
Four PandaSet sequences, frame 40 Deployed bgr vs training-order rgb (019, 056 bike lanes, 065 night)
PandaSet 019, frame 40: p150 vs CPU nuScenes v1.0-mini scene-0103, key-frame 28, CAM_FRONT: non-commercial, CC BY-NC-SA 4.0
KITTI semantics (held out), p150 vs ground truth: non-commercial, CC BY-NC-SA 3.0

PandaSet renders: contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms; resized and annotated with model outputs; Scale AI and Hesai do not endorse this work. The nuScenes render is non-commercial (CC BY-NC-SA 4.0): rendered from the nuScenes dataset, © Motional AD Inc., nuScenes Terms of Use; Motional does not endorse this work. The KITTI render is non-commercial (CC BY-NC-SA 3.0): images and labels from the KITTI Vision Benchmark Suite. Sources, changes and the full attributions: media/ATTRIBUTION.md.

Demo & Performances

Warm, batch 1, 2026-10-09. Latency: the stage bench of OPT_BASELINE.md (code/scripts/bench.py, 100 iterations per stage) on the shipped sample code/tt_sceneseg/samples/highway_normal_1.png (727x414) and on a PandaSet front-camera frame (1920x1080, not shipped), re-checked on the release code; the served rows from uvicorn on the host (the app the container runs) and a loopback client, 50 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load; the device rows repeat to within 0.03 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network (same weights, same pre- and post-processing) on the shipped samples, on 23 public gate frames and on 415 held-out public frames.

Metric Performance
Agreement with the fp32 CPU reference, shipped sample (320x640 class map / published mask) 99.993 % of pixels / 99.994 %; logits PCC 0.9999982
Agreement, 23 public gate frames (9 tutorial images, 2 comma10k, 8 PandaSet, 4 nuScenes) logits PCC ≥ 0.9999952; argmax agreement ≥ 99.943 % (mean 99.983 %)
Agreement, 415 held-out public frames (KITTI 200, comma10k val 100, nuScenes 55, PandaSet 48, comma10k 12) vs ONNX Runtime argmax agreement ≥ 99.752 % (mean 99.980 %, 1st percentile 99.852 %); rgb ≥ 99.812 %
Ground-truth foreground IoU, the 6 shipped comma10k images (training ROI; training data) 0.7288 on p150 vs 0.7287 fp32 CPU (rgb 0.7645 vs 0.7647)
Module PCC vs the fp32 reference (backbone stages, SceneContext, neck, head, logits; each module fed the reference inputs) ≥ 0.99995 (gate 0.999)
Python model() call, shipped sample 727x414 (decode + squash, H2D, trace, D2H, post) 90.85 ms p50 (p99 91.55) · 11.0 frames/s
Python model() call, PandaSet 1920x1080 116.12 ms p50 (p99 119.01) · 8.6 frames/s
Served /predict timing_ms.total (uvicorn on the host, shipped sample; includes the base64 + PNG decode, 16.6 ms) 113.4 ms median (min 112.6)
Served client round trip, loopback (467 kB base64 request) 120.2 ms median
Device trace, one blocking forward 61.14 ms
Back-to-back trace replays 61.06 ms per frame · 16.4 frames/s
Host decode · squash · H2D · D2H · host post (shipped sample) 10.1 · 12.2 · 5.3 · 0.5 · 1.4 ms
from_pretrained load, warm kernel cache 13.1 s

The rgb profile's stage-bench p50s are within 0.35 ms of bgr on every row (the device rows within 0.02 ms); served, its timing_ms.total median was 114.6 ms (20 requests). All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, with the published numerics (two-term bf16 weights, float32 backbone activations as two bf16 terms). Accuracy is agreement with the fp32 CPU reference of the same network; SceneSeg has no paper and no published benchmark, so the dataset metrics below are this port's own measurements. Details: VERIFICATION_2026-10-09.md, OPT_BASELINE.md, OPT_REPORT.md.

Channel order. VisionPilot's deployed C++ pre-processing feeds B, G, R planes (with the ImageNet constants reordered to match) to a network trained on R, G, B: a channel-order bug in the deployment. The default bgr profile reproduces the deployed node; rgb feeds the training order. Both run the same device graph (the order is folded into the first conv's weights) and both are validated against the fp32 CPU reference of their own order. The training order is clearly better, most of all on two-wheelers (foreground IoU unless stated; the datasets are held out except comma10k):

data (ground truth) bgr p150 (fp32 CPU) rgb p150 (fp32 CPU)
comma10k validation split, the authors' own protocol (100 images, training ROI; per-image smoothed IoU) 0.553 (0.553) 0.601 (0.601)
KITTI semantics 2015 (200 frames, 2:1 centre crop; not in the training data; pooled) 0.785 (0.785) 0.814 (0.813)
PandaSet front camera, LiDAR labels projected (48 frames, 434,527 points) 0.821 (0.822) 0.881 (0.881)
PandaSet: share of Bicycle points predicted foreground (1,886 points) 0.056 (0.056) 0.790 (0.791)
nuScenes CAM_FRONT, LiDAR labels projected (55 key-frames) 0.697 (0.697) 0.713 (0.713)
nuScenes: share of motorcycle points predicted foreground (59 points) 0.29 (0.29) 0.71 (0.73)
the 6 shipped comma10k images, training ROI (5 are training images) 0.729 (0.729) 0.764 (0.765)

The p150 column scores the served trace's class maps of all 952 public-frame runs with the same research protocols as the CPU (research/sceneseg/public_data, ss_sanity.py --cls-dir); the dataset labels are mapped to SceneSeg's three classes (vehicles, two-wheelers and people are foreground). The LiDAR rows are a sparse proxy (points 1-40 m projected into the camera), and the comma10k validation split is reconstructed from the upstream code.

Use bgr to reproduce the Autoware VisionPilot output, rgb for better masks.

No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). VisionPilot ran this network as a TensorRT FP16 engine built from SceneSeg_FP32.onnx (autoseg.yaml: backend tensorrt, precision fp16); the upstream model README (Models/model_library/SceneSeg/README.md @ ca58cb50) reports 18.1 FPS FP32 / 26.7 FPS FP16 on an RTX 3060 Mobile (not like-for-like). p150 power was not measured, so no efficiency comparison is made.

Caveats

  • First release: baseline port, optimization pending. The device graph is one metal trace and is kernel-bound, but only 41 % of its 61 ms is conv / matmul compute: the rest is the fp32 glue of the precision policy and data movement. OPT_REPORT.md ranks what comes next.
  • Deployment status in Autoware: SceneSeg is not part of Autoware Universe. It was deployed by the Autoware Foundation's VisionPilot stack (formerly autoware.privately-owned-vehicles) through its generic ROS 2 models package, node scene_seg_model (TensorRT FP16; ONNX Runtime and Zenoh wrappers shared the same pre- and post-processing). Upstream removed that deployment code on 2026-06-15 (last tree 04fa3e80, which this port follows) and the PyTorch model code on 2026-07-06 (last tree ca58cb50); the current VisionPilot 1.0 ships AutoSpeed, AutoSteer and AutoDrive, not SceneSeg. This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions.
  • Channel order: the default bgr reproduces the deployed node, including its channel-order bug (above). rgb is the training order, offered as a load-time option; it is not what VisionPilot ran.
  • Numerics: validated against the fp32 CPU reference of the same graph (ONNX Runtime on the deployed ONNX), not against VisionPilot's TensorRT FP16 engine.
  • Precision policy of this release: every conv, transposed conv and linear runs its weights as two bf16 terms (hi + lo) in separate bias-free convs summed in fp32; the backbone keeps float32 activations and feeds each layer two bf16 terms of them; every bias is added as an fp32 row (a bias fused into a ttnn conv or linear truncates its fp32 result toward zero); HiFi4, fp32 accumulation, packer L1 accumulation; SceneContext and the decoder run bf16 activations. Plain bf16 weights replay in 39.5 ms instead of 61.1 ms, but they missed the 99.5 % agreement gate on a nuScenes frame (0.9949, before fix round 1), and bf16 backbone activations miss it on 3 held-out frames in the CPU emulation (OPT_REPORT.md, "Rejected"). A small residual remains: the stock depthwise and 3x3 convs keep a slight toward-zero bias even without a fused bias, so the hardest held-out frame stays at 99.75 % agreement.
  • Documented deviations from the node: the squash to 640x320 (bit-exact with cv::resize INTER_LINEAR) and the INTER_NEAREST resize back run on the host, the normalisation, the network and the class decision (argmax, lowest index on ties, as the node's strict > scan) on the device; the 3-class map is returned as well; the mask is a PNG (or npz) in JSON, not a ROS image message; inputs are RGB arrays or encoded files (converted to BGR8 inside). SceneSeg's Python tooling also offered a bicubic resize; it is not offered here, because the deployed node never used it. The INT8 QAT ONNX (never deployed) and SceneSegLite (a different 19-class network) are not part of this package.
  • Domain: SceneSeg was trained on ACDC, MUSES, IDDAW, Mapillary Vistas and comma10K (roughly 2:1 views without the ego hood) for a front camera with 52-55° HFOV. On the public frames used here the fp32 CPU reference itself shows the weights' limits, and the p150 output agrees with it: small and distant people and two-wheelers are the weak spot (above all in the deployed bgr order), parking areas and a third of sidewalk points come out as road, trams stay background, and wider cameras (nuScenes 65°) or frames with the ego hood (raw comma10k) score lower.
  • Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication.
  • dispatch="worker" (server: SCENESEG_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (63.82 ms per replay instead of 61.06 ms); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.
  • Batch 1, one frame per request; requests are serialised on the chip. The network input is fixed at 640x320 (the SceneContext reshape hard-codes the 10x20 stride-32 grid); any source image size is accepted, squashed on the host without keeping the aspect ratio, and the mask is resized back with nearest neighbour.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • p150 power was not measured, so no efficiency comparison is made.

Licensing

  • Weights: weights/ of this repo at commit a6829b246a05dba55c356323bbf75f7b1d148499: VisionPilot's SceneSeg_FP32.onnx (Google Drive 1l-dniunvYyFKvLD7k16Png3AsVTuMl9f, sha256 8e509094…c973) and its .pth (Drive 1vCZMdtd8ZbSyHn1LCZrbNKMK7PQvJHxj, sha256 9482f131…89fc) converted to safetensors, re-hosted unchanged with the upstream Apache-2.0 LICENSE, NOTICE and provenance.json (the upstream README covers the model weights with Apache-2.0; the card AutowareFoundation/SceneSeg holds no weights). The upstream model card lists the training data: ACDC, MUSES, IDDAW, Mapillary Vistas and comma10K (BDD100K held out); check those datasets' terms before commercial use.
  • Pre- and post-processing ported from autowarefoundation/vision_pilot VisionPilot/middleware_recipes/ROS2/models and middleware_recipes/common/backends @ 04fa3e80 (Apache-2.0; removed upstream on 2026-06-15).
  • Port and serving code (code/): Apache-2.0. patches/tt-metal-eth-dispatch.patch modifies tt-metal (Apache-2.0).
  • Sample data (only redistributable data ships; the public-dataset frames of the accuracy tables are not in this repository):
    • code/tt_sceneseg/samples/*.png: the 9 images of the upstream SceneSeg tutorial (autowarefoundation/vision_pilot @ ca58cb50, Models/tutorials/assets/images/), Apache-2.0 (samples/LICENSE-vision_pilot.txt); the original photo source is not stated upstream.
    • code/tt_sceneseg/samples/comma10k/: six comma10k images with their masks (commaai/comma10k @ 6c205fe), MIT, © 2020 Comma.ai, Inc. (samples/comma10k/LICENSE-comma10k.txt); five of them are SceneSeg training images.
  • Demo media (media/, sources and changes in media/ATTRIBUTION.md):
    • PandaSet renders: Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: resized, annotated with model outputs. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021.
    • nuScenes render (media/sceneseg_nuscenes_scene0103_k28_front_tt_NC.jpg), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020.
    • KITTI render (media/sceneseg_kitti_heldout_tt_NC.jpg), non-commercial, CC BY-NC-SA 3.0: Contains images and labels from the KITTI Vision Benchmark Suite, semantic segmentation benchmark 2015 (A. Geiger, P. Lenz, C. Stiller, R. Urtasun; H. Alhaija et al.), https://www.cvlibs.net/datasets/kitti/. Non-commercial use only; adaptations under the same license. Changes: cropped to 2:1, resized, prediction and ground truth overlaid. Cite: A. Geiger et al., CVPR 2012; H. Alhaija et al., IJCV 2018; M. Menze and A. Geiger, CVPR 2015.
    • Tutorial-image renders: Apache-2.0; comma10k renders: MIT.

Provenance

These are the exact sources the container image was built from:

component built from
tt-metal 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (sha256 08d0ddf6…; dirty tree: the image includes the patch)
weights changh95/sceneseg-p150@a6829b246a05dba55c356323bbf75f7b1d148499, files weights/SceneSeg_FP32.onnx, weights/provenance.json, weights/LICENSE, weights/NOTICE (ONNX sha256 8e509094…c973, checked at load)
Autoware reference autowarefoundation/vision_pilot 04fa3e80d0b0a9d2491c07991a9cb48101b2c2cb (deployment, VisionPilot/middleware_recipes); model code ca58cb50
shared package ttaw 0.22.0, vendored as code/tt_sceneseg/ttaw from the Autoware ports' shared common repository at commit 2218be3 (code/tt_sceneseg/ttaw/VENDORED.json: version, commit and per-file sha256)
code/ digest (image) 7bbcad160628eb66 (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json)
image tt-model/sceneseg-p150:f070810e48b2 (sha256:f070810e48b2785605cef8dc572be0ddfa4db141ce47c93c71266a0b581f9a18)
base images build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json)
built 2026-10-09T18:40:28+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/sceneseg-p150

Quantized
(1)
this model

Collection including changh95/sceneseg-p150