frnet-p150
FRNet (Autoware lidar_frnet): the LiDAR semantic segmentation network of Autoware's opt-in autoware_lidar_frnet node, ported to one Tenstorrent Blackhole p150 with tt-nn. One LiDAR sweep in the sensor frame (x, y, z, raw 0-255 intensity) in; a class id and probability for every point (27 T4 classes, Autoware's interpolated points included) and the node's filtered-cloud mask out, with the node's exact pre- and post-processing. One image, two serve profiles: ot128 (Hesai OT128, the default) and qt128 (Hesai QT128).
Weights: AutowareFoundation/lidar_frnet v2.0 · Paper: arXiv:2312.04484 · Autoware package: autoware_lidar_frnet · Training code: tier4/AWML (projects/FRNet) · Port: code/
Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Precision: fp32 weights; the point-wise layers (point MLPs, fusions, head) take fp32 activations as two bf16 terms with fp32 accumulation; the frustum CNN runs fp32 feature maps with bf16 conv operands; HiFi4 math with fp32 accumulation everywhere; no bfp8. Both serve profiles run in this configuration, and all numbers on this card were measured in it.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/frnet-p150 --exclude "image/*" --local-dir frnet-p150 && cd frnet-p150
pip install -e . # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub; ttnn and torch come from tt-metal
pip install -e ".[server,test]" # optional: the HTTP server and the tests
Run the snippet from the model repo root: the sample and its calibration preset are paths relative to it.
from tt_frnet import FRNet
with FRNet.from_pretrained(device_id=0) as model: # variant="qt128" for the Hesai QT128 profile
out = model("code/tt_frnet/samples/synthetic_ot128.npz", # one sweep in the sensor frame: path, (N, 4) array, .npy/.npz/.pcd, or a PointCloud
calibration={"preset": "synthetic_ot128"}) # its sensor -> base_link transform, for the node's ego crop box
print(out.class_counts()) # points per class (raw + interpolated)
labels, scores, xyz = out.label_ids, out.scores, out.points[:, :3]
from_pretraineddownloadsAutowareFoundation/lidar_frnetat the pinned commita53a1c11b0b(tagv2.0, 85 MB) to your HF cache, checks the ONNX sha256, opens the chip, builds the graph of the variant and captures the metal trace. The first load compiles the kernels (156.6 s for ot128 and 151.9 s for qt128 with an empty JIT cache); later loads take 7.5 / 6.6 s.- The trace is captured during the load, so the first call compiles nothing: it took 320 ms (ot128) / 241 ms (qt128), against 305 / 215 ms for the second call.
- The
withblock releases the trace and closes the chip. Withoutwith, callmodel.close(). - For your own cloud, pass its transform:
calibration={"T_base_link_from_sensor": T}(4×4, or{x, y, z, roll, pitch, yaw}), or{"ego_crop_box": None}to disable the ego crop box.
| Input | ONE LiDAR sweep in the sensor frame (x forward, y left, z up; Autoware's ~/input/pointcloud): path (.bin / .npy / .npz / .pcd), raw bytes with fmt=, an (N, C) float array or tensor, or a PointCloud; fields x, y, z, intensity (raw, 0-255). calibration: {"T_base_link_from_sensor": ...} or {"preset": name} for the node's ego crop box, optional ego_crop_box (bounds, or None to disable it). |
| Options | class_probability_threshold=0.05, filter_classes="drivable_surface" (Autoware's filter: they change filter_keep only). from_pretrained(device_id=0, variant="ot128" | "qt128", dispatch="eth", weights_dir=None, device=None). |
| Output | FRNetSegmentation: label_ids uint8 [N] (Autoware's class_id), scores float32 [N] (its probability), points float32 [N, 4], filter_keep bool [N] (the filtered cloud), num_points_raw, input_index [N_raw], class_names (27), meta (crop, counts, skipped), timing_ms. Order: the input points kept by the ego crop box (input order), then Autoware's interpolated points. |
| Methods | out.to_dict() gives the /predict JSON. out.class_counts() gives the points per class, out.to_dicts() one row per class. tt_frnet.viz.render_bev(points, label_ids) draws a bird's-eye view. |
- The API gives the same output as the HTTP server
/predict: both share the decoders, the device trace and the host post-processing (checked on the device bytest_api_equals_server, both profiles). variantselects the network at load time. Each is tied to its sensor's FOV and frustum grid:ot128+15° / −25°, 128 × 1024 (Autoware's defaultsensor_model);qt128±52.6°, 128 × 256.- A frame outside the variant's profile (5,000-160,000 / 120,000 points, 3,000-60,000 cells) returns an empty output with the node's message in
meta["skipped"], as the node publishes nothing; nothing runs on the device. - One model uses one chip; calls from several threads are serialised.
- Full reference:
code/PYTHON.md. Runnable example:examples/quickstart.py(also writesquickstart_bev.png, the p150 classes from above next to the stored CPU reference).
Serving (HTTP)
tt-model pull changh95/frnet-p150 --with-weights
tt-model serve changh95/frnet-p150 # QT128: tt-model serve --profile qt128 changh95/frnet-p150 (options go before the repo id); tt-cli: tt serve changh95/frnet-p150
python3 code/tt_frnet/server/client.py --points code/tt_frnet/samples/synthetic_ot128.npz --calib-preset synthetic_ot128 --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/frnet-p150
- The image does not contain the weights.
--with-weightsputs them in your HF cache. - The server uses port 20000 (or the next free port). It is ready when the log shows
Application startup complete. - Two serve profiles (
tt-model profiles changh95/frnet-p150):ot128(the default) andqt128. Each serves one variant. client.pybuilds the request with the standard library only; add--url http://127.0.0.1:20000to send it,--calib my_calib.jsonfor your own transform.POST /predict:points(base64.bin/.npy/.npz/.pcd, sensor frame, fields x, y, z, intensity);calibration(T_base_link_from_sensor, or{"preset": name}; optionalego_crop_box,nulldisables the box); optionalparams(class_probability_threshold,filter_classes),output_format. AlsoGET /health,GET /info(with the effective default of every knob and the calibration presets),GET /v1/models(stub). Contract:SERVING.mdsection 3.
{"model": "frnet-p150", "frame_id": "sensor", "num_points": 52613, "num_points_raw": 52613,
"labels": {"format": "npz", "key": "labels", "dtype": "uint8", "shape": [52613], "data": "..."},
"scores": {"format": "npz", "key": "scores", "dtype": "float32", "shape": [52613], "data": "..."},
"points": {"format": "npz", "key": "points", "dtype": "float32", "shape": [52613, 4], "data": "..."},
"filter_keep": {"format": "npz", "key": "filter_keep", "dtype": "bool", "shape": [52613], "data": "..."},
"input_index": {"format": "npz", "key": "input_index", "dtype": "int32", "shape": [52613], "data": "..."},
"class_names": ["drivable_surface", "other_flat_surface", "sidewalk", "manmade", "vegetation", "car", "..."],
"class_counts": {"drivable_surface": 2913, "other_flat_surface": 20118, "sidewalk": 5, "manmade": 22898, "vegetation": 6526, "car": 13, "semi_trailer": 140, "...": 0},
"meta": {"variant": "ot128", "num_points_input": 53411, "cropped_by_ego_box": 798, "num_points_interpolated": 0, "num_unique_coors": 52613,
"ego_crop": {"enabled": true, "bounds": [-1.53, -1.15, -0.1, 6.1, 1.15, 3.1], "reference_frame": "base_link", "...": "..."},
"plan": {"n": 52613, "m": 52613, "n_cap": 160000, "m_cap": 60000, "...": "..."},
"filter": {"classes": ["drivable_surface"], "class_probability_threshold": 0.05, "kept": 44747}, "...": "..."},
"timing_ms": {"decode": 24.711, "preprocess": 47.664, "device": 229.164, "postprocess": 5.018, "model_call": 281.879, "total": 310.689}}
labels/scores/points/filter_keep/input_indexare base64 NPZ arrays (tt_frnet.io.decode_array, or numpy on the decoded bytes);class_countsandmetaare plain JSON.- The first
num_points_rawpoints are input points (input_indexmaps them back to the input rows); the rest are the points Autoware's interpolation appends, published as Autoware does.
Demo
The p150 output on the shipped samples: synthetic scans generated by this repository (a ray cast against boxes; no dataset involved). Left: the p150 classes; middle: the stored fp32 CPU reference; right: the points whose class differs.
On public driving datasets (p150 outputs, OT128 profile; the frames themselves are not in this repository). The caption of each image gives its agreement with the fp32 CPU reference on that frame.
PandaSet renders: contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms; converted to the sensor frame, rendered as bird's-eye views and downscaled camera images with projected points; Scale AI and Hesai do not endorse this work. The nuScenes renders are non-commercial (CC BY-NC-SA 4.0): rendered from the nuScenes dataset, © Motional AD Inc., nuScenes Terms of Use; Motional does not endorse this work. Sources, changes and the full attributions: media/ATTRIBUTION.md.
Demo & Performances
Warm, batch 1, 2026-10-09. Latency: the stage bench of OPT_BASELINE.md (code/scripts/bench.py, 100 iterations per stage; re-checked on the release commit to within 0.05 ms on the device rows) on the shipped synthetic sample (OT128: 53,411 points, 52,613 after the ego crop; QT128: 45,783 points, 69,553 after Autoware's interpolation) and on a PandaSet frame (107,146 points; not shipped); the served rows from uvicorn on the host (the app the container runs) and a loopback client, 50 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load; the device rows repeat to 0.01 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network (ONNX Runtime on the deployed ONNX, with the port's Autoware pre- and post-processing, which is bit-identical to the research reference), on every raw point the node scores (public frames); on the shipped samples, on every output point, Autoware's interpolated points included.
| Metric | Performance |
|---|---|
| Class agreement with the fp32 CPU reference, shipped sample | OT128 99.55 % of 52,613 points (filtered-cloud mask 99.77 %) · QT128 99.62 % of 69,553 (99.87 %) |
| Class agreement, all 113 converted public frames per profile: PandaSet 019 / 029 (32 frames, every 5th) and nuScenes v1.0-mini mini_val (81 key-frames), production API | OT128: PandaSet 99.66 % pooled (worst frame 99.35 %), nuScenes 99.19 % (worst 98.37 %) · QT128: PandaSet 99.29 % (98.69 %), nuScenes 99.44 % (98.10 %) |
Frozen device gates, 17 OT128 / 5 QT128 frames (PandaSet, nuScenes, Autoware's sample rosbag; gates pred_probs PCC ≥ 0.998, agreement ≥ 98.5 %) |
all pass: OT128 min PCC 0.99969, min agreement 98.907 % (nuScenes 0916 kf40, the binding frame); QT128 0.99980 / 99.158 % |
| Module PCC vs the fp32 reference (replay outputs, PandaSet 019 f40) | every gated module ≥ 0.9996: worst OT128 point_fusion1 0.999729, QT128 attention4 0.999649 (gate 0.999) |
| Mapped-label mIoU, OT128 (a sanity check, not a paper metric: dataset labels mapped to the 27 T4 classes, see Caveats) | PandaSet 32 frames: coarse p150 0.714 · CPU 0.715, fine (classes with ≥ 1,000 labelled points) 0.447 · 0.448 · nuScenes 81 frames: coarse p150 0.483 · CPU 0.483, fine (same rule) 0.451 · 0.453 |
Python model() call, shipped sample (host pre-processing, H2D, trace, D2H, host post-processing) |
OT128 292.6 ms p50 (3.4 calls/s) · QT128 221.1 ms p50 (4.5 calls/s) |
Python model() call, PandaSet frame (107,146 points; not shipped) |
OT128 355.7 ms · QT128 244.3 ms p50 |
Served /predict timing_ms.total (uvicorn on the host, the shipped sample) |
OT128 310.4 ms · QT128 228.9 ms median (including 24.3 / 20.7 ms of request decode) |
| Served client round trip, loopback | OT128 328.1 ms · QT128 247.3 ms median |
| Device trace, one blocking forward (shipped sample) | OT128 210.02 ms · QT128 133.53 ms |
| Back-to-back trace replays | OT128 209.95 ms per forward (4.76 frames/s; 211.64 ms on PandaSet) · QT128 133.43 ms (7.49 frames/s; 134.87 ms) |
| Host pre-processing · H2D · D2H · un-permute + post-processing (shipped sample) | OT128 55.3 · 5.42 · 7.65 · 7.1 + 5.0 ms; QT128 54.2 · 4.04 · 5.80 · 6.1 + 6.8 ms |
from_pretrained load: empty JIT cache / warm cache |
OT128 156.6 s / 7.5 s · QT128 151.9 s / 6.6 s |
All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150 at 1350 MHz AICLK, with the published precision (FRNET_POINT_TERMS=2, FRNET_MAP_DTYPE=float32, FRNET_CONV_TERMS=1). The wider 113-frame agreement set is not gated: 3 of its 226 frame-profile pairs fall below 98.5 % (nuScenes scene-0916 key-frames 11 and 1 on OT128, 98.37 % and 98.499 %; key-frame 0 on QT128, 98.10 %), all on the out-of-domain Singapore scene (see Caveats). The mIoU rows are indicative only: approximate label mappings, small sets, sensors and taxonomy different from the training data, and no published number exists for these weights. Details: VERIFICATION_2026-10-09.md, OPT_BASELINE.md, OPT_REPORT.md.
No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). Autoware runs this network as a TensorRT engine (fp16 by default, per the weights card); this port was validated against the fp32 ONNX Runtime reference, not against a TensorRT engine. p150 power was not measured, so no efficiency comparison is made.
Caveats
- First release: baseline port, optimization pending. Each profile runs as one metal trace of 1,104 (ot128) / 759 (qt128) programs and is kernel-bound (op-to-op gaps 0.3 %). 63 % (ot128) / 75 % (qt128) of the device time is the point-wise layers kept at higher precision, whose cost is DRAM traffic of fp32 partial sums, not compute; every point-wise op also runs on the profile's full capacity (160,000 / 120,000 rows) whatever the frame.
OPT_REPORT.mdranks what comes next. - Deployment status in Autoware: FRNet is an opt-in, standalone node.
autoware_lidar_frnetships in autoware_universe with its own launch file (lidar_frnet.launch.xml,sensor_model:=ot128 | qt128), but no Autoware launch configuration wires it into the default perception stack. This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions. - Narrow accuracy margin on one frame. On the hardest public frame of the release (nuScenes scene-0916 key-frame 40, OT128: a 32-beam sensor, out of the training domain, low-confidence outputs), the p150 agrees with the fp32 reference on 98.907 % of the points, 0.41 points above the 98.5 % gate. Every cheaper precision configuration measured during the port failed that frame (bf16 maps 98.28 %, an fp32 × fp32 point island 97.50 %), which is why the release keeps the slower two-term point layers and fp32 maps. Beyond the gated frames, 3 of the 226 frame-profile pairs of the wider public set fall just below 98.5 % (98.10-98.499 %), all on the same Singapore scene-0916. It is the binding constraint, and an explicit precision item, of the optimization phase.
- Points near a class boundary can take the other class: the device's probabilities differ slightly from the fp32 reference (
pred_probsPCC ≥ 0.9997 on every public frame), and the argmax turns a small difference into a different class. The figures are in the agreement rows above. - Documented deviations from the node:
- Autoware's CUDA races are replaced by fixed, documented policies: interpolated points are appended in row-major pixel order after the raw points (Autoware's atomic append order varies), the fp16 interpolation grid keeps the farthest point with ties to the later point, a duplicated cell keeps the last voxel (as ONNX Runtime); the output is deterministic;
- the projection uses numpy's float32
arctan2/arcsin, which may differ from CUDA'satan2f/asinfby an ulp, so a point exactly on a bin edge could change cell; - the pre-processing, the argmax and the class filter run on the host (bit-identical to the CPU reference); the network on the device;
- inputs are arrays or files, not
PointCloud2messages; the filtered cloud is returned as a mask (filter_keep), not converted to Autoware'soutput_format; - label names 22 / 23: Autoware's
class_namesreaddebris/stroller, but the weights were trained with stroller = 22 and debris = 23 (AWML config). The port keeps Autoware's names for parity.
- Precision policy of this release: fp32 weights; the point-wise linears split their fp32 inputs and weights into two bf16 terms (three exact bf16 products, fp32 accumulation; the device's own fp32 matmul is TF32-like and fails the gates); fp32 feature maps in the frustum CNN with bf16 conv operands and fp32 weights; HiFi4 with fp32 accumulation and packer L1 accumulation; K1
segment_reduce(a customgeneric_opkernel) is bit-exact. Pinned inserve.env(FRNET_POINT_TERMS=2,FRNET_MAP_DTYPE=float32,FRNET_CONV_TERMS=1). - Domain: the weights were trained by TIER IV on about 16,000 frames of the T4 dataset from Hesai OT128 / QT128 sensors (not public), with a 27-class T4 taxonomy. No public dataset has these sensors or labels, so the card shows agreement with the fp32 reference, not accuracy. On the public frames the reference itself shows the gap: PandaSet's Hesai Pandar64 shares the OT128's vertical FOV and segments cleanly (coarse mapped mIoU 0.714 on the p150, 0.715 for the reference, OT128), nuScenes' 32-beam HDL-32E much less so (0.483 for both), and the QT128 network on these sensors is out of distribution (coarse 0.19 on PandaSet; used here only as a numerical test). Traffic cones are almost never found and pedestrians next to structures often become
manmade, in the reference as on the p150. - Validation scope: agreement with the fp32 CPU reference (ONNX Runtime on the deployed ONNX, with the port's bit-exact Autoware pre- and post-processing) on public driving data (PandaSet, nuScenes v1.0-mini) and Autoware's sample rosbag, plus a mapped-label mIoU sanity check (not a paper metric: different taxonomy and sensors). The knobs were validated at their Autoware defaults.
- Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (
patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication. dispatch="worker"(server:FRNET_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (211.56 ms per ot128 replay instead of 209.95 ms; 153.71 ms instead of 133.44 ms for qt128); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.- Batch 1, one sweep per request; requests are serialised on the chip; one variant per process (the serve profile). The device input has a fixed size per profile (160,000 / 120,000 points, 60,000 cells), whatever the cloud.
- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150 power was not measured, so no efficiency comparison is made.
Licensing
- Weights: AutowareFoundation/lidar_frnet at tag
v2.0(commita53a1c11b0b28d31fd50a03abef5db0e830a5b3c), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card's training data: the T4Dataset (TIER IV), about 16,000 frames (4,000 frames × 4 surrounding sensors), trained with AWML; it states no further data licence notice. - Pre- and post-processing ported from autoware_universe
perception/autoware_lidar_frnet(Apache-2.0); the FOV, grids, class names, palette and profiles are read from the weights repo'sml_package_frnet_*.param.yamlat load time. - Port and serving code (
code/): Apache-2.0.patches/tt-metal-eth-dispatch.patchmodifies tt-metal (Apache-2.0). - Sample data:
code/tt_frnet/samples/synthetic_ot128.npzandsynthetic_qt128.npzare synthetic scans generated bycode/scripts/make_sample.py(no dataset involved), with their calibration presets and stored CPU-reference outputs: Apache-2.0. Only these samples ship; the public-dataset frames of the accuracy rows are not in this repository. - Demo media (
media/, sources and changes inmedia/ATTRIBUTION.md):- PandaSet renders: Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: the Pandar64 sweep re-expressed in an estimated sensor frame, bird's-eye views rendered, the front-camera image downscaled and dimmed, points and labels overlaid. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021.
- nuScenes renders (
media/*_NC.png), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020; W. K. Fong et al., Panoptic nuScenes, RA-L 2022 (lidarseg labels). frnet_synthetic_*renders: Apache-2.0.- The agreement and mapped-mIoU rows use PandaSet (CC BY 4.0) and nuScenes v1.0-mini with nuScenes-lidarseg (CC BY-NC-SA 4.0); only the resulting numbers are on this card.
Provenance
These are the exact sources the container image was built from:
| component | built from |
|---|---|
| tt-metal | 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (sha256 08d0ddf6…; dirty tree: the image includes the patch) |
| weights | AutowareFoundation/lidar_frnet@a53a1c11b0b28d31fd50a03abef5db0e830a5b3c (tag v2.0), files frnet_*.onnx, ml_package_frnet_*.param.yaml, deploy_metadata.yaml (ONNX sha256 ot128 d29bf595…, qt128 647af927…) |
| Autoware reference | autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (perception/autoware_lidar_frnet, package 0.53.0) |
| shared package | ttaw 0.20.0, vendored as code/tt_frnet/ttaw from the Autoware ports' shared common repository at commit 89dec49 (code/tt_frnet/ttaw/VENDORED.json: version, commit and per-file sha256); it holds the K1 segment_reduce kernel |
code/ digest (image) |
f0b62790446a51a8 (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json) |
| image | tt-model/frnet-p150:85cbf6cc72b9 (sha256:85cbf6cc72b9c535c2f44cec2b2145c8fb940d214874c23c918adfc66a51769c) |
| base images | build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json) |
| built | 2026-10-09T17:04:34+00:00 by tt-model 0.1.0 |
Model tree for changh95/frnet-p150
Base model
AutowareFoundation/lidar_frnet






