diffusion-planner-p150
Diffusion Planner v5.0 (Autoware diffusion_planner): the network Autoware deploys in autoware_diffusion_planner, ported to one Tenstorrent Blackhole p150 with tt-nn. The whole plan runs on the chip as one metal trace: the scene encoder, the 11 DiT decoder evaluations of the DPM-Solver++(2M) loop with their solver updates, and the turn-indicator head. The Autoware planner tensors in (ego and neighbour histories, lanes, route, polygons, line strings, goal, ego shape, turn-indicator history); an 8 s ego trajectory, the predicted 8 s paths of the neighbours and a turn-indicator command out, with the node's exact pre- and post-processing.
Weights: AutowareFoundation/diffusion_planner v5.0 · Paper: arXiv:2501.15564 · Autoware package: autoware_diffusion_planner · Training code: tier4/Diffusion-Planner (TIER IV training fork) · Port: code/
Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Numerics (the default and the serve.env pins): fp32 residual streams and solver state, HiFi4 with fp32 accumulation, split hi / lo matmuls for the mixer inputs and every decoder linear, fp32 LayerNorm in the mixers and the decoder, fp32 matmul attention in the fusion encoder and the decoder: the configuration the end-to-end accuracy gates need (Caveats). All numbers on this card were measured in this configuration.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/diffusion-planner-p150 --exclude "image/*" --local-dir diffusion-planner-p150 && cd diffusion-planner-p150
pip install -e . # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub; ttnn and torch come from tt-metal
pip install -e ".[server,test]" # optional: the HTTP server and the tests
Run the snippet from the model repo root: code/tt_diffusion_planner/samples/kashiwanoha_dense.npz is a path relative to it.
from tt_diffusion_planner import DiffusionPlanner
with DiffusionPlanner.from_pretrained(device_id=0) as model: # weights -> your HF cache, trace captured
out = model(inputs="code/tt_diffusion_planner/samples/kashiwanoha_dense.npz") # the 15 raw planner tensors: .npz path, its bytes, or {name: array}
print(out.columns) # x, y, yaw, cos, sin, velocity, acceleration (base_link, 0.1-8.0 s)
print(out.poses[:5])
print(out.turn_indicator["command_name"], out.predicted_agents.shape)
from_pretraineddownloads the three v5.0 ONNX files anddiffusion_planner.param.json(58.9 MB) ofAutowareFoundation/diffusion_plannerat the pinned commit423efde67f5(tagv5.0) to your HF cache (no token needed), checks their sha256, opens the chip, builds the graph and captures the metal trace. The first load compiles the kernels (315 s with an empty JIT cache, firmware and every kernel compiled); later loads take about 9 s (device open, reading and uploading the weights, warm-up and trace capture).- The trace is captured during the load, so the first call is as fast as the later calls and no call compiles anything.
- The
withblock releases the trace and closes the chip. Withoutwith, callmodel.close().
| Input | inputs=: the 15 raw tensors of the node's create_input_data() (ego frame, before normalization; batch 1): an .npz path or its bytes, a {name: array} mapping, or the /predict JSON envelope. Checked against DiffusionPlanner.INPUT_SCHEMA (names, shapes, finite values). Converting ROS messages and the Lanelet2 map into these tensors stays with the client. |
| Options | velocity_smoothing_window=8, stopping_threshold=0.3, turn_indicator_keep_offset=-1.25, return_denoising_steps=False. from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None). |
| Output | Trajectory: poses float32 [80, 7] (x, y, yaw, cos, sin, velocity, acceleration; base_link), turn_indicator (command, logits, probabilities), predicted_agents float32 [N, 80, 5], timing_ms, meta (the neighbour rows, force stop, the solver iterates on request). |
| Methods | out.to_dict() gives the /predict JSON. out.to_dicts() gives one dict per trajectory point. |
- The API gives the same output as the HTTP server
/predict: both share the decoders, the device trace and the host post-processing (checked on the device bytest_api_equals_server). - The API is stateless: one call is one plan. What the Autoware node keeps between plans stays with the caller: the turn-indicator hold window (1.0 s; the node's manager ships as
tt_diffusion_planner.host.postprocess.TurnIndicatorManager), the initial solver statesampled_trajectories(zeros = the node's default temperature 0; noise for a temperature > 0; the previous plan for the RTC prefix) and the agent and ego histories. Details:code/PYTHON.md"What the caller keeps". - One model uses one chip; calls from several threads are serialised.
- Full reference:
code/PYTHON.md. Runnable example:examples/quickstart.py(also writesquickstart_bev.png, the input tensors and the plan from above).
Serving (HTTP)
tt-model pull changh95/diffusion-planner-p150 --with-weights
tt-model serve changh95/diffusion-planner-p150 # or with tt-cli: tt serve changh95/diffusion-planner-p150
python3 code/tt_diffusion_planner/server/client.py --inputs code/tt_diffusion_planner/samples/kashiwanoha_dense.npz --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/diffusion-planner-p150
- The image does not contain the weights.
--with-weightsputs them in your HF cache. - The server uses port 20000 (or the next free port). It is ready when the log shows
Application startup complete. - One serve profile, the default (
tt-model profiles changh95/diffusion-planner-p150).serve.envpins ETH dispatch, 1 CQ and the numerics knobs. - Run the request lines from the model repo root (the
hf downloadof the Quickstart): the client and the sample are files of this repo.code/tt_diffusion_planner/server/client.pyneeds only the Python standard library (no numpy); add--url http://127.0.0.1:20000to send the request. POST /predict:inputs(the 15 raw tensors of the Autoware node'screate_input_data()in the ego frame, before normalization: a base64.npzor{"format": "json", "arrays": {...}}); optionalparams(velocity_smoothing_window8,stopping_threshold0.3,turn_indicator_keep_offset-1.25,return_denoising_stepsfalse),output_format. AlsoGET /health,GET /info,GET /v1/models(stub). Contract:SERVING.mdsection 3.
The response for the shipped sample (served on the p150; trajectory cut to 3 of 80 rows):
{
"model": "diffusion-planner-p150",
"frame_id": "base_link",
"meta": {"predicted_agent_columns": ["x", "y", "yaw", "cos", "sin"], "predicted_agent_rows": [0, 1, 2, "... 85 more"], "force_stop": false, "time_from_start_s": [0.1, 0.2, 0.3, "... 77 more"], "valid_counts": {"ego": 1, "neighbor": 88, "static": 0, "lane": 123, "route": 17, "polygon": 0, "line_string": 60, "goal": 1, "ego_shape": 1, "turn": 1}},
"timing_ms": {"preprocess": 3.8, "device": 105.3, "postprocess": 3.2, "total": 122.8, "decode": 8.5, "model_call": 112.5},
"num_poses": 80,
"columns": ["x", "y", "yaw", "cos", "sin", "velocity", "acceleration"],
"trajectory": [[0.365, 0.01, -0.017, 0.9962, -0.0169, 3.8059, -0.0431], [0.7631, -0.0022, -0.0414, 0.9966, -0.0413, 3.8016, -0.5485], [1.1524, -0.0282, -0.0675, 0.9942, -0.0674, 3.7468, -0.4708], "... 77 more rows"],
"turn_indicator": {"command": 1, "command_name": "DISABLE", "keep_selected": true, "held": false, "logits": [-15.960084915161133, -5.125061511993408, -5.230083465576172, -0.7658059597015381, 5.504596710205078], "probabilities": [1.651601411190029e-09, 8.38487030705437e-05, 7.548934809165075e-05, 0.006556871347129345, 0.993184506893158]},
"predicted_agents": {"format": "npz", "key": "predicted_agents", "dtype": "float32", "shape": [88, 80, 5], "data": "<base64 npz>"}
}
trajectory: 80 points at 0.1-8.0 s inbase_link(the ego frame of the input tensors), post-processed like the node's~/output/trajectory: velocity from consecutive points, forward moving average overvelocity_smoothing_windowpoints, force stop (poses frozen once the smoothed speed falls belowstopping_thresholdwhile the ego moves), acceleration by finite difference.yawis whattf2::getYawreads from the node's (unnormalised) quaternion;cos/sinare the raw network outputs.turn_indicator.command: 0 NO_COMMAND, 1 DISABLE, 2 ENABLE_LEFT, 3 ENABLE_RIGHT (KEEP repeats the last input report); the node's 1 s hold window needs state across calls and is not applied (see above).predicted_agents: one 80-point path (x, y, yaw, cos, sin) per non-empty neighbour row, in input order (meta.predicted_agent_rows).
Demo
On nuScenes v1.0-mini planning instants (p150 outputs; the scenes are converted from the dataset into the planner tensors and are not in this repository). Non-commercial, CC BY-NC-SA 4.0. The grey path with hollow dots is the logged drive; the model never saw nuScenes (see Caveats).
nuScenes renders: rendered from the nuScenes dataset (v1.0-mini, CAN bus expansion and map expansion v1.3), © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use; non-commercial use only; Motional does not endorse this work. The kashiwanoha sample is derived from the Apache-2.0 map AutowareFoundation/map-carla-kashiwanoha. Sources, changes and the full attributions: media/ATTRIBUTION.md.
Demo & Performances
Warm, batch 1; the stage bench 2026-10-08, the served rows, load times and the re-check 2026-10-09. Latency: the stage bench of OPT_BASELINE.md (code/scripts/bench.py, 100 iterations per stage) on the shipped sample kashiwanoha_dense.npz (88 neighbours, 123 lanes, 17 route lanes, 60 line strings; the device time is the same for every scene: every plan computes the full capacities); the served rows from uvicorn on the host (the app the container runs, with its serve pins) and a loopback client, 50 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load by several ms; the device rows repeat to ±0.05 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network on 99 scenes: the 2 shipped samples, 5 research scenes and 92 nuScenes v1.0-mini planning instants (the frozen end-to-end gates), and against the research pipeline's ONNX Runtime outputs as an independent oracle.
| Metric | Performance |
|---|---|
Agreement with the fp32 CPU reference, shipped sample kashiwanoha_dense (8 s ego plan, 88 neighbours) |
ego max 2.3 cm / mean 1.1 cm; turn command identical; neighbours: median per-agent max 8.6 cm |
Agreement, shipped sample straight_road |
ego max 4.9 cm / mean 1.8 cm; turn command identical; neighbours 6.3 cm |
| Agreement, all 99 gated scenes (2 samples, 5 research scenes, 92 nuScenes-mini instants) | worst ego max 0.313 m / mean 0.143 m (gates 1.0 / 0.3 m; nuscenes/scene-0103_kf14); ego mean: median 1.3 cm, 95th percentile 6.2 cm; turn command identical 99 / 99; neighbours: median per-agent max ≤ 0.086 m (gate 1.5 m) |
| Agreement with an independent oracle: the research pipeline's ONNX Runtime outputs (raw x0), 92 nuScenes instants + the 33-plan scene-0061 sequence | worst ego max 0.313 m / mean 0.143 m, turn command identical 92 / 92; sequence: worst ego max 0.118 m / mean 0.049 m, turn 33 / 33; every instant within the gates |
| Module PCC vs the fp32 reference (encoder categories, encoding, teacher-forced decoder evaluation; gate 0.999) | ≥ 0.999952 / 0.999978 / ≥ 0.9999995 |
| Open-loop vs the nuScenes log, 92 instants (a sanity check against one human driver, not a planning metric) | p150 ADE / FDE at 8 s 4.75 / 11.87 m (CPU reference 4.74 / 11.87; constant velocity 4.99 / 13.06); turn command = logged blinker 88 / 92 (CPU 88 / 92) |
Python model() call, shipped sample (host pre-processing, H2D, trace, D2H, host post-processing) |
117.9 ms p50 (p99 134.6) · 8.5 plans/s |
Served /predict timing_ms.total (uvicorn on the host, the shipped sample) |
123.1 ms median (min 122.5; of which decode 8.6) |
Served client round trip, loopback (base64 .npz request, 0.15 MB) |
127.5 ms median |
| Device trace, one blocking plan (encoder + 11 DiT evaluations + 10 solver updates + turn head) | 102.13 ms |
| Back-to-back trace replays | 102.04 ms per plan · 9.80 plans/s |
| Host pre-processing · pack · host tensors · H2D · D2H · host post-processing | 4.75 · 0.45 · 2.41 · 1.14 · 0.58 · 4.32 ms |
from_pretrained load: empty JIT cache / warm cache |
315 s / 8.6 s |
All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, with the pinned numerics of serve.env. Accuracy is agreement with the fp32 CPU reference of the same Autoware network (same weights, same pre- and post-processing); no dataset-level accuracy is claimed: the paper's benchmark is nuPlan closed loop (an account-gated dataset and simulator, not run here), and the deployed v5.0 weights were trained by TIER IV on data that is not public, so no public benchmark is in-domain. Details: VERIFICATION_2026-10-08.md, OPT_BASELINE.md, OPT_REPORT.md.
No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). Autoware's CHANGELOG quotes 5.13 ms mean (300 runs) for an older single-step engine of this planner on an RTX PRO 6000 Blackwell with TensorRT (precision not stated); it is not like-for-like with this v5.0 multi-step port. p150 power was not measured, so no efficiency comparison is made.
Caveats
- First release: baseline port, optimization pending. The plan is one metal trace and kernel-bound (6,282 programs; op-to-op gaps 4.4 ms of a 103.7 ms span). The first optimization target is the precision cost: the numerics defaults that the end-to-end gates need (split hi / lo matmuls, fp32 LayerNorm, fp32 matmul attention) cost 33.4 ms per plan (102.0 ms vs 68.6 ms with the first device round's defaults, which fail the gates); fused kernels are to recover it without dropping precision. A second sink of similar size is unrelated to precision: the encoder's channel-MLP and pre-projection matmuls run on 4-8 cores (25.8 ms).
OPT_REPORT.mdranks what comes next. - Deployment status in Autoware:
autoware_diffusion_planneris an alternative to the default rule-based planning stack, selected withplanning_setting:=diffusion_planner(package README); it is aimed at Autoware's proposed new planning framework. This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions or closed-loop vehicle control. - Stateless API: the node's state between plans (the turn-indicator hold window, the RTC prefix and temperature of the initial solver state, the agent buffers and the ego history) is the client's ("What the caller keeps" in
code/PYTHON.md). The node's guidance services (start / stop / centerline guidance) are off, as in the node's default. - Precision policy of this release: fp32 residual streams and solver state; HiFi4 with fp32 accumulation for every matmul; the ego / neighbour pre-projection as a pad-relative fp32 island; split hi / lo matmuls (bf16 hi + fp32 lo parts, ~1e-5 relative, because a device fp32 matmul rounds its operands like TF32) for the mixer inputs and every decoder linear; an fp32 LayerNorm decomposition in the mixers and the decoder (the fused
ttnn.layer_normloses the per-entity signal on the mixers' offset-dominated rows); fp32 matmul attention in the fusion encoder and the decoder; the turn head in fp32; the other weights and the hidden MLP activations of the mixers and the fusion encoder in bf16. The plan's sensitivity to these choices is chaotic per scene: with the first device round's numerics, one nuScenes instant (scene-0103_kf14) moved 0.35 m on average while the module PCCs differed only in the 5th decimal. So every numerics change is re-checked on all 99 scenes. - Validation scope: the p150 output agrees with the fp32 CPU reference on 99 scenes (above). The two shipped samples and the five research scenes are synthetic-but-faithful scenes built with a Python port of the node's tensor construction; the 92 nuScenes instants are converted from nuScenes v1.0-mini, a domain the model never saw (Singapore and Boston, right-hand traffic in Boston, oracle tracks from 2 Hz annotations, no traffic-light states, no speed limits, stop areas instead of stop lines). On them the plans are plausible but conservative (moving plans about 17 % shorter than the logged drive); the open-loop numbers above are a sanity check against one human driver, not a planning metric.
- Neighbour predictions: the gate is the median over agents of each agent's max displacement, so single agents can differ more. On the shipped sample
kashiwanoha_dense(container smoke, served p150 output vs the stored CPU reference) the per-agent max displacement is 8.6 cm median and 0.26 m at the 90th percentile; the worst of the 88 agents differs by 4.31 m. The ego plan is gated on its max and mean; neighbour paths only on that median. - Documented deviations from the node: the output stays in
base_link(the node transforms it tomapwith the ego pose);yawis whattf2::getYawreads from the node's quaternion of the unnormalised cos / sin rotation (it differs from atan2(sin, cos) when |(cos, sin)| ≠ 1, as in the node); the speed masks follow the node's TensorRT path (> FLT_EPSILON); the turn indicator is decided without the hold window;delayis accepted and ignored (the node's multi-step mode never reads it). - Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (
patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication. dispatch="worker"(server:DIFFUSION_PLANNER_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (101.87 ms per replay, the same as ETH within 0.2 %); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.- Batch 1, one plan per request; requests are serialised on the chip. The shapes of the v5.0 export (320 neighbours, 140 lanes, 25 route lanes, 10 polygons, 60 line strings, 31 history and 80 future steps) and 10 DPM-Solver steps are compiled into the trace, and every plan computes them in full, so the device time does not depend on the scene.
- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150 power was not measured, so no efficiency comparison is made.
Licensing
- Weights: AutowareFoundation/diffusion_planner at tag
v5.0(commit423efde67f5414734da43a7ad856c17ceb8b51aa), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card states that TIER IV trained the models on TIER IV synthetic and real driving data; the dataset composition is not publicly documented. - Pre- and post-processing ported from autoware_universe
planning/autoware_diffusion_planner(Apache-2.0). - Port and serving code (
code/): Apache-2.0.patches/tt-metal-eth-dispatch.patchmodifies tt-metal (Apache-2.0). - Sample data:
code/tt_diffusion_planner/samples/kashiwanoha_dense.npzis derived from the Apache-2.0 Lanelet2 map AutowareFoundation/map-carla-kashiwanoha (0.2.0) with a scripted ego, route and agents;straight_road.npzis a procedural scene generated by this repo. Both Apache-2.0, each with its stored CPU-reference output (*.reference.json). Only these redistributable samples ship; the nuScenes-derived planning instants of the accuracy tables are not in this repository. - Demo media (
media/, sources and changes inmedia/ATTRIBUTION.md):- nuScenes renders (
media/*_NC.*), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset (v1.0-mini, CAN bus expansion and map expansion v1.3), © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020. - Renders of the shipped samples: Apache-2.0 (kashiwanoha map: AutowareFoundation/map-carla-kashiwanoha@0.2.0, Apache-2.0).
- nuScenes renders (
Provenance
These are the exact sources the container image was built from:
| component | built from |
|---|---|
| tt-metal | 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (sha256 08d0ddf6…45cc, 4 files; dirty tree: the image includes the patch) |
| weights | AutowareFoundation/diffusion_planner@423efde67f5414734da43a7ad856c17ceb8b51aa (tag v5.0), files diffusion_planner_encoder.onnx, diffusion_planner_decoder.onnx, diffusion_planner_turn_indicator.onnx, diffusion_planner.param.json (sha256 2856886a…ca49, eb30c0c0…57ca, 07acfb58…a732, ee3145b6…a268, checked at load) |
| Autoware reference | autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (planning/autoware_diffusion_planner, package 0.53.0, multi_step mode) |
| shared package | ttaw 0.20.0, vendored as code/tt_diffusion_planner/ttaw from the Autoware ports' shared common repository at commit 89dec49 (code/tt_diffusion_planner/ttaw/VENDORED.json: version, commit and per-file sha256) |
code/ digest (image) |
c0e7abb7888098a9 (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json) |
| image | tt-model/diffusion-planner-p150:3b96d8ea7190 (sha256:3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909) |
| base images | build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json) |
| built | 2026-10-09T04:37:19+00:00 by tt-model 0.1.0 |
Model tree for changh95/diffusion-planner-p150
Base model
AutowareFoundation/diffusion_planner










