yolox-p150

YOLOX-s+ seg16 (Autoware tensorrt_yolox): the network Autoware deploys in autoware_tensorrt_yolox, its camera 2D object detector, ported to one Tenstorrent Blackhole p150 with tt-nn. One camera image of any size in (converted to BGR8, as the node receives it); 2D boxes of 8 classes (UNKNOWN, CAR, TRUCK, BUS, BICYCLE, MOTORBIKE, PEDESTRIAN, ANIMAL), the node's Autoware ROIs and a 16-class semantic mask out, with the node's exact pre- and post-processing. Weights: AutowareFoundation/tensorrt_yolox v1.0 · Paper: arXiv:2107.08430 · Autoware package: autoware_tensorrt_yolox · Training code: tier4/trt-yoloXP (TIER IV YOLOX-opt; the export fork tier4/YOLOX is not public) · Port: code/

Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Numerics clip6 (the default and only serve profile): bf16 activations, conv weights and biases as two bf16 terms (hi + lo) with fp32 accumulation. All numbers on this card were measured in this configuration.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.

hf download changh95/yolox-p150 --exclude "image/*" --local-dir yolox-p150 && cd yolox-p150
pip install -e .                        # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub; ttnn and torch come from tt-metal
pip install -e ".[server,test]"         # optional: the HTTP server and the tests

Run the snippet from the model repo root: code/tt_yolox/samples/test_image.jpg is a path relative to it.

from tt_yolox import YOLOX

with YOLOX.from_pretrained(device_id=0) as model:      # weights -> your HF cache, traces captured
    out = model("code/tt_yolox/samples/test_image.jpg")   # path, bytes, PIL image or RGB uint8 HxWx3 array of any size

for d in out.to_dicts()[:5]:
    print(d["label"], d["score"], d["box_xyxy"])
mask = out.extras["semseg"].mask              # uint8 class ids (semseg_color_map.csv), e.g. 620x960
  • from_pretrained downloads the seg16 ONNX, label.txt and semseg_color_map.csv (62 MB) of AutowareFoundation/tensorrt_yolox at the pinned commit 424d6b9cb94 (tag v1.0) to your HF cache, opens the chip, builds the graph and captures the metal trace. The first load compiles the kernels (about 179 s with an empty JIT cache, firmware and every kernel compiled); later loads take about 10 s (device open, weight preparation and trace capture).
  • The trace is captured during the load, so no call compiles anything. The host letterbox builds a lookup table once per source image size, so the first call of a new size is a little slower (172 vs 157 ms for the second call on the shipped sample).
  • The with block releases the trace and closes the chip. Without with, call model.close().
Input One camera image: path, PNG / JPEG bytes, PIL image, RGB uint8 HxWx3 array (any size), or images=[...] like the server. Converted to BGR8 and letterboxed to 960x960 with Autoware's exact resize. An OpenCV / ROS bgr8 array must be flipped first (bgr[:, :, ::-1]).
Options score_threshold=0.35, nms_threshold=0.7, is_roi_overlap_segmentation=True, overlap_roi_score_threshold=0.3, is_publish_color_mask=False. from_pretrained(device_id=0, variant="clip6" | "relu6" | "relu", dispatch="eth", weights_dir=None, device=None).
Output Detections2D: boxes_xyxy float32 [N, 4] (original pixels), scores [N], label_ids [N] (label.txt order), labels, extras["semseg"] (Mask2D, uint8 class ids), extras["rois"] (Autoware ROI records), timing_ms. Sorted by score.
Methods out.to_dict() gives the /predict JSON. out.to_dicts() gives the detection list; out.extras["semseg"].mask the class-id mask.
  • The API gives the same output as the HTTP server /predict: both share the decoders, the device trace and the host post-processing (checked on the device by test_api_equals_server).
  • variant selects the numerics at load time: clip6 (default) reproduces the saturation of Autoware's int8 engine (clip_value: 6.0), relu6 is the training semantics and relu the ONNX file as shipped. All three pass the same accuracy gates; the server reads YOLOX_VARIANT.
  • One model uses one chip; calls from several threads are serialised.
  • Full reference: code/PYTHON.md. Runnable example: examples/quickstart.py (also writes quickstart.png, the mask and the boxes over the image).

Serving (HTTP)

tt-model pull  changh95/yolox-p150 --with-weights
tt-model serve changh95/yolox-p150       # or with tt-cli: tt serve changh95/yolox-p150
printf '{"images": [{"camera": "CAM_FRONT", "data": "%s"}]}' "$(base64 -w0 code/tt_yolox/samples/test_image.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/yolox-p150
  • The image does not contain the weights. --with-weights puts them in your HF cache.
  • The server uses port 20000 (or the next free port). It is ready when the log shows Application startup complete.
  • One serve profile, clip6 (the default; tt-model profiles changh95/yolox-p150).
  • python3 code/tt_yolox/server/client.py --image CAM_FRONT=<image> --out req.json builds the same request (standard library only); add --url http://127.0.0.1:20000 to send it.
  • POST /predict: images (one entry: base64 JPEG / PNG of the camera image, any size, decoded to BGR8 like Autoware), optional params (score_threshold 0.35, nms_threshold 0.7, is_roi_overlap_segmentation true, overlap_roi_score_threshold 0.3), output_format. Also GET /health, GET /info, GET /v1/models (stub). Contract: SERVING.md section 3.
{"model": "yolox-p150", "frame_id": "camera", "num_detections": 9,
 "detections": [
   {"label": "CAR", "label_id": 1, "score": 0.9356, "box_xyxy": [344.0, 376.0, 398.0, 410.0]},
   {"label": "TRUCK", "label_id": 2, "score": 0.9307, "box_xyxy": [556.0, 308.0, 832.0, 465.0]}],
 "semseg": {"format": "png", "key": "mask", "dtype": "uint8", "shape": [620, 960], "data": "<base64 PNG>"},
 "rois": [{"x_offset": 344, "y_offset": 376, "width": 54, "height": 34, "label": "CAR", "label_id": 1, "autoware_label": 1, "score": 0.935619}],
 "meta": {"variant": "clip6", "image_hw": [827, 1280], "mask_hw": [620, 960], "scale": 0.75},
 "timing_ms": {"preprocess": 63.1, "device": 77.7, "postprocess": 6.2, "total": 175.0, "decode": 18.3, "model_call": 146.9}}
  • box_xyxy are original-image pixels: the integer Autoware ROI (x_offset, y_offset, width, height, also returned in rois with the remapped Autoware ObjectClassification id autoware_label).
  • semseg is the Autoware mask: the argmax over the 16 classes of semseg_color_map.csv, at network resolution and cropped to the un-letterboxed region (620x960 for a 1280x827 image), with the ROI overlay of the node (is_roi_overlap_segmentation).

Demo

Input (code/tt_yolox/samples/test_image.jpg, the shipped sample) Detections and mask on p150 (media/yolox_autoware_test_image_tt.jpg)

The p150 output next to the fp32 CPU reference on the same image; the bottom panel marks in white the pixels where the two published masks differ (0.13 % of them):

On public driving datasets (p150 outputs; the frames themselves are not in this repository). The caption of each image gives its agreement with the fp32 CPU reference.

PandaSet 005, frame 40, front camera PandaSet 005, frames 20-49 (3 s at 10 Hz), front camera
PandaSet 005, frame 40, all six cameras PandaSet 019, frame 40, all six cameras
PandaSet 065 at night, frame 40, front camera PandaSet 005, frame 40, left camera: p150 vs CPU (the lowest mask agreement of the 24 gate frames)
nuScenes v1.0-mini scene-0103, key-frame 28, CAM_FRONT: non-commercial, CC BY-NC-SA 4.0

PandaSet renders: contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms; resized and annotated with model outputs; Scale AI and Hesai do not endorse this work. The nuScenes render is non-commercial (CC BY-NC-SA 4.0): rendered from the nuScenes dataset, © Motional AD Inc., nuScenes Terms of Use; Motional does not endorse this work. Sources, changes and the full attributions: media/ATTRIBUTION.md.

Demo & Performances

Warm, batch 1, 2026-10-08. Latency: the stage bench of OPT_BASELINE.md (code/scripts/bench.py, 100 iterations per stage) on the shipped sample code/tt_yolox/samples/test_image.jpg (1280x827) and on a PandaSet front-camera frame (1920x1080, not shipped); the served rows from uvicorn on the host (the app the container runs) and a loopback client, 50 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load; the device rows repeat to 0.001 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network on the shipped sample, on the 24 public frames of the device gates and on 265 public frames in all (PandaSet day and night, nuScenes v1.0-mini day and night).

Metric Performance
Agreement with the fp32 CPU reference, shipped sample 9 of 9 detections matched one-to-one (same label, IoU ≥ 0.5; min IoU 0.986, max |Δscore| 0.0024); published mask 99.87 % of pixels equal; det-output PCC 0.99956, seg-logit PCC 0.99999
Agreement, the 24 public gate frames (PandaSet 005 / 019 frame 40 and nuScenes mini_val scene-0103 / 0916, six cameras each) det-output PCC ≥ 0.99885; mask agreement ≥ 99.42 % (mean 99.76 %); no detection with score ≥ 0.40 unmatched (324 / 329 one-to-one)
Agreement, 265 public frames (PandaSet 005 / 019 / 065-night, nuScenes mini_val + night scenes, the demo frames) 3,370 of 3,394 CPU detections matched one-to-one (99.3 %); one object with score ≥ 0.40 changes class (PEDESTRIAN 0.461 on the CPU, MOTORBIKE 0.456 on the p150: a near-tie); published mask agreement ≥ 98.42 % on that frame, ≥ 99.42 % on every other (mean 99.77 %)
2D recall vs projected 3D boxes, 238 public images (a sanity check, not a paper metric) IoU ≥ 0.3: cars 448 / 462 (CPU 448 / 462), pedestrians 421 / 571 (CPU 421 / 571)
Module PCC vs the fp32 reference (stem, dark2-5, PAN outputs, head levels; each module fed the reference inputs) ≥ 0.99994 (gate 0.999)
Python model() call, shipped sample 1280x827 (host decode + letterbox, H2D, trace, D2H, host post) 154.4 ms p50 (p99 163.6) · 6.5 frames/s; 177.2 ms p50 in a re-check on a busier host
Python model() call, PandaSet 1920x1080 200.0 ms p50 (p99 219.9) · 5.0 frames/s
Served /predict timing_ms.total (uvicorn on the host, shipped sample) 172.6 ms median (min 171.1)
Served client round trip, loopback (base64 JPEG request) 178.8 ms median
Device trace, one blocking forward 51.49 ms
Back-to-back trace replays 51.43 ms per forward · 19.45 frames/s
Host decode · letterbox · H2D · D2H · host post (shipped sample) 18.5 · 51.7 · 23.2 · 2.3 · 6.1 ms
from_pretrained load: empty JIT cache / warm cache 179 s / 9.8 s

All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, clip6 numerics with two-term bf16 weights (the serve pins). Accuracy is agreement with the fp32 CPU reference of the same Autoware network (same weights, same pre- and post-processing); no dataset-level accuracy is claimed, because none exists for these weights (their training data and labels are TIER IV's, and the 16 segmentation classes have no public ground truth). Details: VERIFICATION_2026-10-08.md, OPT_BASELINE.md, OPT_REPORT.md.

No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). Autoware runs this network as a TensorRT int8 engine with every layer output clipped to 6 (clip_value: 6.0); the default clip6 numerics of this port reproduce that saturation deterministically, without the int8 quantization noise. p150 power was not measured, so no efficiency comparison is made.

Caveats

  • First release: baseline port, optimization pending. The device graph is one metal trace and is kernel-bound, but 82 % of its 51.4 ms is layout changes, data movement and precision glue; the host letterbox takes 50-70 ms per frame. OPT_REPORT.md ranks what comes next.
  • Deployment status in Autoware: this is the default model of the autoware_tensorrt_yolox node, but Autoware's default top-level launch (autoware.launch.xml) starts no YOLOX node: its perception component is LiDAR-only, and the camera 2D detector runs only in the camera-LiDAR(-radar) object-recognition launch with enable_2d_detection:=true (default false). This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions.
  • Numerics: the default clip6 mode reproduces the saturation of Autoware's int8 engine, not its int8 quantization: the port was validated against the fp32 CPU reference of the same graph, not against a TensorRT engine. The ONNX file as shipped (relu) differs strongly from clip6 on public frames (median mask agreement 68 % on PandaSet, CPU reference).
  • Precision policy of this release: bf16 activations; the conv weights and biases as two bf16 terms (hi + lo) summed in fp32; HiFi4, fp32 accumulation, packer L1 accumulation; the head predictions in fp32. Plain bf16 weights halve the trace time (26.3 ms) but fail the public-frame mask gate (97.55 % on PandaSet 005 left camera, gate 99 %), because the bf16 rounding of the weights and biases dominates the error. bf16 still moves single scores slightly: an object near the 0.35 score threshold, a near-duplicate pair near the 0.7 NMS IoU, or an object whose two best classes score almost the same can flip. On 265 public frames this happened to one object above 0.40: a person on a two-wheeler that the CPU reference calls PEDESTRIAN (0.461) and the p150 MOTORBIKE (0.456) on nuScenes scene-0103 key-frame 32 (details in the verification report).
  • Documented deviations from the node: the mask is the argmax of the logits (the softmax is dropped; identical masks on all 75 sample × mode goldens); the decoder assumes the float overload of exp (the double one gives identical boxes on all 75 goldens) and the letterbox the float fma (the double one would change 0-2 of 3,000 sampled pixels by 1); which overloads Autoware's build picks is UNVERIFIED; the letterbox and the post-processing run on the host (the letterbox bit-exact with the node's CUDA kernel), the class decision (argmax) on the device; the mask is returned as a PNG (or npz), not as the node's run-length-encoded image; inputs are RGB arrays or encoded files (converted to BGR8 inside), not ROS image messages.
  • Domain: the upstream card does not document the training data of these T4 weights. On the public frames used here (PandaSet, nuScenes: other cameras, undistorted images, left-hand traffic in Singapore) the fp32 CPU reference itself shows the weights' limits, and the p150 output agrees with it ("Demo & Performances"): SF buses often come out as TRUCK, sidewalk is practically never predicted (it comes out as road), crosswalk_others fires on road texture, and construction barriers get no box.
  • Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication.
  • dispatch="worker" (server: YOLOX_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (53.94 ms per replay instead of 51.43 ms); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.
  • Batch 1, one frame per request; requests are serialised on the chip. The network input is fixed at 960x960 (Autoware letterbox); any source image size is accepted, and the letterbox runs on the host.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • p150 power was not measured, so no efficiency comparison is made.

Licensing

  • Weights: AutowareFoundation/tensorrt_yolox at tag v1.0 (commit 424d6b9cb94c939ae9ce6a22380e46c13f76b1da), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card states that the training datasets, schedules and evaluation metrics of its T4-finetuned variants are not publicly documented; the seg16 model is TIER IV's, trained on pseudo-labelled TIER IV (T4) data, as its file name says.
  • Pre- and post-processing ported from autoware_universe perception/autoware_tensorrt_yolox (Apache-2.0); the label, remap and colour tables in code/tt_yolox/host/data/ are copies of the package's and the weights repo's files.
  • Port and serving code (code/): Apache-2.0. patches/tt-metal-eth-dispatch.patch modifies tt-metal (Apache-2.0).
  • Sample data: code/tt_yolox/samples/test_image.jpg is the Autoware package's own test image (autoware_universe perception/autoware_tensorrt_yolox/test/test_image.jpg, Apache-2.0). Only this redistributable sample ships; the public-dataset frames of the accuracy tables are not in this repository.
  • Demo media (media/, sources and changes in media/ATTRIBUTION.md):
    • PandaSet renders: Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: resized, annotated with model outputs. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021.
    • nuScenes render (media/yolox_nuscenes_scene0103_k28_front_tt_NC.jpg), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020.
    • Autoware test image renders: Apache-2.0.

Provenance

These are the exact sources the container image was built from:

component built from
tt-metal 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (sha256 08d0ddf6…; dirty tree: the image includes the patch)
weights AutowareFoundation/tensorrt_yolox@424d6b9cb94c939ae9ce6a22380e46c13f76b1da (tag v1.0), files yolox-sPlus-opt-pseudoV2-T4-960x960-T4-seg16cls.onnx, label.txt, semseg_color_map.csv (ONNX sha256 73b38124…16a8, checked at load)
Autoware reference autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (perception/autoware_tensorrt_yolox, package 0.53.0)
shared package ttaw 0.20.0, vendored as code/tt_yolox/ttaw from the Autoware ports' shared common repository at commit 89dec49 (code/tt_yolox/ttaw/VENDORED.json: version, commit and per-file sha256)
code/ digest (image) 6fb069cffd2b7ba3 (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json)
image tt-model/yolox-p150:3ec596ba8f97 (sha256:3ec596ba8f979a68d041cb8d55ea7d8c01c92687d3daa9332ab3e350f5f030f3)
base images build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json)
built 2026-10-08T07:07:20+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/yolox-p150

Finetuned
(1)
this model

Collection including changh95/yolox-p150

Paper for changh95/yolox-p150