|
Download code/tt_diffusion_planner/ttaw/API.md from changh95/diffusion-planner-p150: direct link, hf CLI and curl.
- Browser
- Download file 147 kB
-
https://huggingface.co/changh95/diffusion-planner-p150/resolve/main/code/tt_diffusion_planner/ttaw/API.md
- Command line
-
hf download hf://changh95/diffusion-planner-p150/code/tt_diffusion_planner/ttaw/API.md
-
curl -L -o API.md https://huggingface.co/changh95/diffusion-planner-p150/resolve/main/code/tt_diffusion_planner/ttaw/API.md
147 kB
| # ttaw API guide (v0.23.0) | |
| `ttaw` is the shared code of the 13 Autoware ports to one Blackhole p150b (PLAN.md section 1). Source of truth: | |
| `common/ttaw/`. Every bundle gets a vendored copy at `bundles/<model>-p150/code/<pkg>/ttaw/` and imports it | |
| relatively (`from .ttaw.trace import TraceRunner`). Tests: `common/tests/host` (CPU, fake ttnn) and | |
| `common/tests/device` (p150, through `bin/devrun`). Changes: `common/CHANGELOG.md`. | |
| Rules that keep one tree valid top-level and vendored: | |
| - relative imports only; importing any module never imports ttnn / torch / onnx, opens no device, reads no | |
| environment and touches no network (host tests and the image's `verify:` step import everything without a device); | |
| - every optimization knob is read once at model build and has an env A/B switch (`ttaw.knobs`); | |
| - kernel `.cpp` files are package data under `ops/kernels/`, located with `ops.kernel_path(name)`. | |
| Contents: 1. bundle wiring - 2. device - 3. TraceRunner - 4. tensors - 5. weights - 6. precision and knobs - | |
| 7. metrics, goldens, gates - 8. profiling - 9. I/O and outputs - 10. ModelBase - 11. server - 12. vendoring - | |
| 13. measured facts and pitfalls - 14. image pre-processing (C14) - 15. LiDAR host pipeline (C10-C12, C15, C16) - | |
| 16. CNN ops: conv builders (C17), up-sampling and interpolation (C18) - 17. attention (C20) - | |
| 18. LiDAR device modules: gather-form scatter (C19), SECOND + SECONDFPN (C24), CenterHead (C25) - | |
| 19. grid_sample helpers (C23) - 20. ResNet builders (C27) - 21. query heads: top-k (C21), heatmap peaks (C22), | |
| TransFusion query head (C26) - 22. pillar feature net and input staging (C24 companion) - 23. segment reductions (K1) - | |
| 24. sparse-conv rulebooks (C13) - 25. gather-GEMM sparse encoder (C28). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 1. Wiring a bundle (thin wrappers over ttaw) | |
| ```bash | |
| python common/tools/vendor.py bundles/<model>-p150 # committed ttaw/ (HEAD) -> code/<pkg>/ttaw + VENDORED.json | |
| python common/tools/vendor.py bundles/<model>-p150 --check # drift report (exit 1 on drift; section 12) | |
| ``` | |
| `research/BUNDLE_TEMPLATE` is the canonical wiring (`instantiate_bundle.py --out` renders it and vendors ttaw; the | |
| values file's `INPUT_KIND` picks its flavour: `lidar` as below, `camera`, `multi_camera`, `lidar_camera`, `planner` | |
| with their own `api.py` / `io.py` / stub / smoke / tests, `research/BUNDLE_TEMPLATE/TEMPLATE_NOTES.md` "Flavours"): | |
| `device.py` and `io.py` bind ttaw to the model, `api.py` subclasses `ModelBase`, `server/app.py` calls `create_app`, | |
| and the stdlib `server/client.py` / `server/smoke_test.py` run `ttaw/server/{client,smoke}.py` by path: | |
| ```python | |
| # code/<pkg>/device.py: ttaw.device bound to the port's validated open parameters (the model class uses the same dict) | |
| from .ttaw.device import DeviceConfig, close_device, describe_device # noqa: F401 | |
| DEVICE_DEFAULTS = {"num_command_queues": 1, "l1_small_size": 32768, "trace_region_size": 64 << 20} | |
| def device_config(**overrides): ... # DEVICE_DEFAULTS < <ENV>_* / TT_DEVICE_ID < non-None overrides | |
| def open_device(device_id=None, *, dispatch=None, **overrides): ... # device_config(...).open() | |
| # code/<pkg>/io.py: ttaw.io re-exported, decode_points / load_points rebound to the model's layout | |
| from .ttaw.io import * # noqa: F401,F403 | |
| DEFAULT_POINT_FIELDS = ("x", "y", "z", "intensity") | |
| # code/<pkg>/api.py (section 10 has the full hook contract) | |
| from .ttaw.api_base import ModelBase | |
| class CenterPoint(ModelBase): | |
| MODEL_NAME, ENV_PREFIX = "centerpoint-p150", "CENTERPOINT" | |
| DEVICE_DEFAULTS = device.DEVICE_DEFAULTS | |
| ... | |
| # code/<pkg>/server/app.py (section 11) | |
| from pathlib import Path | |
| from ..api import CenterPoint | |
| from ..ttaw.server.app import ServerSpec, create_app, parse_mesh_shape # noqa: F401 (tt-model.yaml verify: line) | |
| app = create_app(ServerSpec(model_name=CenterPoint.MODEL_NAME, env_prefix="CENTERPOINT", model_cls=CenterPoint, | |
| task="LiDAR 3D object detection", default_weights=CenterPoint.DEFAULT_REPO, | |
| calib_dir=Path(__file__).resolve().parents[1] / "calib")) | |
| ``` | |
| `pyproject.toml` package data: `"<pkg>.ttaw" = ["API.md", "VENDORED.json"]`, `"<pkg>.ttaw.ops" = ["kernels/*.cpp", | |
| "kernels/*.hpp", "kernels/*.h"]`. The tt-model `extra_code` path `<pkg>` already ships the sub-package. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 2. Device (C01, `ttaw.device`) | |
| ```python | |
| from .ttaw.device import DeviceConfig, describe_device, device_session, open_device, compute_grid | |
| with device_session(dispatch="eth", num_command_queues=2, trace_region_size=64 << 20) as dev: | |
| print(describe_device(dev)) # {'dispatch': 'eth', 'grid': '12x10', 'cores': 120, 'num_command_queues': 2, | |
| # 'arch': 'blackhole', 'fallback': None, 'eth_patch': True, ...} | |
| gx, gy = compute_grid(dev) # never hard-code 12x10: program configs read the grid | |
| cfg = DeviceConfig.from_env("CENTERPOINT", num_command_queues=1) # <ENV>_DISPATCH / _NUM_CQS / _L1_SMALL / | |
| dev = cfg.open() # _TRACE_REGION / _WORKER_L1_SIZE, TT_DEVICE_ID | |
| ``` | |
| | name | signature / meaning | | |
| |---|---| | |
| | `open_device` | `(device_id=0, *, dispatch="eth", num_command_queues=1, l1_small_size=32768, trace_region_size=64<<20, worker_l1_size=None, allow_fallback=True)` -> ttnn device. `dispatch`: `eth` (12x10), `worker` (11x10, A/B only), `auto` (ETH if the patch marker is in `tt_metal/impl/dispatch/topology.cpp`, else WORKER + `RuntimeWarning`). A failed ETH open warns and falls back to WORKER; `allow_fallback=False` re-raises (use it in container smoke / CI). Warns if ETH does not give 12x10 | | |
| | `DeviceConfig` | frozen dataclass of the arguments above; `.from_env(prefix, env=None, **defaults)`, `.open()` | | |
| | `describe_device` | `(device, device_id=None)` -> dict: `dispatch` (used), `dispatch_requested`, `fallback`, `num_command_queues`, `grid` "12x10", `grid_x`, `grid_y`, `cores`, `arch`, open sizes, `eth_patch`. `dispatch="unknown"` for devices not opened by ttaw | | |
| | `open_info(device)` | what `open_device` recorded for this device object (`{}` for a device opened elsewhere; the record is tied to the object, never to a reused `id`) | | |
| | `close_device` / `device_session` | close (syncs first; idempotent: a second close of the same device is a no-op) / context manager that always closes | | |
| | `resolve_dispatch(dispatch, *, home=None)`, `eth_dispatch_patch_present(home=None)`, `tt_metal_home()` | patch detection (file read; `TT_METAL_HOME`, `TT_METAL_RUNTIME_ROOT`, or the tree holding `ttnn`, found without importing it) | | |
| | `compute_grid(dev)` / `core_grid(dev)` / `full_core_range_set(dev)` | `(x, y)` / `ttnn.CoreGrid` / one-rectangle `ttnn.CoreRangeSet` of the whole grid | | |
| | constants | `DISPATCH_MODES`, `PATCH_MARKER`, `DEFAULT_L1_SMALL_SIZE`, `DEFAULT_TRACE_REGION_SIZE`, `P150_ETH_GRID` | | |
| `worker_l1_size` is absolute bytes of allocatable L1 per core; tt-metal's default is computed at open time | |
| (about 1,461,248 B on this tree, TT_PLATFORM.md section 1). Shrinking it grows the kernel-config ring buffer. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 3. Traces (C02, `ttaw.trace`) | |
| Every device stage runs inside traces. `TraceRunner` owns the persistent tensors, the warm-up, the captures, the | |
| replays and the readback. | |
| ```python | |
| import ttnn | |
| from .ttaw.trace import TraceRunner, pack_outputs | |
| runner = TraceRunner(device, num_command_queues=2, name="centerpoint") # CQ count defaults to the open's | |
| runner.add_input("pillars", shape=(1, 1, 40000, 64), dtype="bfloat16") # ROW_MAJOR in DRAM by default | |
| runner.add_input("index", init=warm_idx, dtype="uint32") # give real data if the graph gathers | |
| runner.add_param("score_thr", 0.35) # RT-dev: fp32 [1,1,1,1] TILE | |
| runner.add_state("prev_bev", shape=(1, 1, 22500, 256), dtype="bfloat16", pingpong=True) # temporal models | |
| def forward(ctx): # ctx[name]: input / param / state read buffer; no host I/O in here | |
| x = model.backbone(ttnn.to_layout(ctx["pillars"], ttnn.TILE_LAYOUT)) | |
| bev = model.temporal(x, ctx["prev_bev"]) | |
| ctx.write_state("prev_bev", bev) # ttnn.copy into the state's write buffer (part of the trace) | |
| heat = ttnn.gt(model.head(bev), ctx["score_thr"]) | |
| return pack_outputs({"heat": heat, "boxes": model.box(bev)}) # one D2H for several outputs | |
| runner.add_variant("default", forward) # one trace per variant (buckets: "n8k", "n16k", ...) | |
| runner.capture() # warm-up of every variant, then strict captures | |
| out = runner("default", inputs={"pillars": p, "index": idx}, params={"score_thr": 0.4}) # {"heat": ..., "boxes": ...} | |
| runner.release() # or `with TraceRunner(...) as runner:` | |
| ``` | |
| **What it guarantees.** Persistent tensors exist (with defined contents) before the first capture; every variant is | |
| warmed `warmup_runs` times before *any* capture, and so are the runner's own eager programs (the `stage_inputs` | |
| staging copy, the state <-> bank copies) and the callables registered with `add_eager_warmup(fn)`; capture runs with | |
| `device.set_program_cache_misses_allowed(False)` | |
| (a miss raises `... program cache miss occurred, but cache misses are forbidden` naming the op), `end_trace_capture` | |
| runs in `finally` and a failed capture is released, and the program-cache size must not change during capture; | |
| adding a variant after a capture releases all traces, warms the new one and recaptures everything; warm-up writes to | |
| states are undone (`reset_state`) before capture; unchanged params are not re-uploaded. | |
| **No program may be compiled after the first capture.** A new program's kernel binaries are a DRAM buffer allocated | |
| when it is first enqueued, in the address space of the traces' freed intermediates, so a replay can overwrite them | |
| (tt-metal `tech_reports/AdvancedPerformanceOptimizationsForModels/TraceCorrectness.md`: corruption or a hang). Any | |
| eager op the model runs between replays (host-fallback glue, an eager `ttnn.to_layout`, `tensors.upload_u8`, | |
| `fp32_island`, ...) must therefore run once before the first capture: register it with `add_eager_warmup(fn)`. The | |
| runner refuses its own eager work that compiles after a capture (`stage_fn` copies, `save_state` / `load_state`, | |
| `reset_state` from a device tensor of another spec, `run_eager`) with `RuntimeError: ... compiled a new program after | |
| capture`, and every port runs its device tests once with `TT_METAL_TRACE_ALLOC_TRACKING=1` (the tracker refuses a | |
| replay while such a buffer is alive; ttaw 0.2.0's `stage_inputs` path tripped it: | |
| `logs/ttaw/review_old_staged_tracking_probe.log`). | |
| **1CQ / 2CQ.** 1CQ: uploads and replays on CQ0. 2CQ: CQ1 waits for the last replay's event, uploads, records; CQ0 | |
| waits, replays, records (the `check_dispatch.py` pattern; the events exist even after a partially failed `capture()` | |
| left some variants runnable). `stage_inputs=True` (2CQ only) is the `tt_cnn` executor | |
| pattern: CQ1 writes DRAM staging copies while the previous replay still runs, an eager `ttnn.copy` (or your | |
| `add_input(..., stage_fn=lambda staging, dst: ...)`, e.g. a reshard into a sharded L1 input) moves them into the | |
| trace inputs on CQ0; that copy is compiled in the first `capture()` (a staging buffer mirrors every | |
| `write_input`). `read(variant, cq_id=1)` (2CQ) waits for that trace's completion event on the host and reads on CQ1, | |
| so CQ0 can already run the next segment. | |
| **Segments.** Split a long graph into variants run back to back; hand intermediates from segment k to k+1 through an | |
| in-place state (`ctx.write_state("mid", t)` in seg1, `ctx["mid"]` in seg2): a variant cannot see another variant's | |
| outputs at warm-up time, a state buffer exists from the start. | |
| **Ping-pong state.** `pingpong=True` allocates two buffers and captures two traces per variant (phase 0 reads A / | |
| writes B, phase 1 the reverse). A variant that writes the ping-pong states must write all of them ("stepping"); the | |
| phase flips after each run of a stepping variant; read-only variants follow the current phase. In-place states | |
| (default) read and write one buffer, so read what you need before the write. Reset with `reset_state(name, value)` | |
| (host data, or a device tensor copied on the device). `write_state` traces one `ttnn.copy`; compute the new state | |
| straight into `ctx.write_target(name)` (`ttnn.add(old, x, output_tensor=ctx.write_target("mem"))`, then | |
| `ctx.write_state("mem", new)`) and no copy is traced (probe P13: one program less per state per frame). | |
| `write_state` / `write_output` check shape, layout (ttnn.copy keeps it) and dtype changes (TILE only) before tracing. | |
| For a ring buffer use `ttnn.copy` / slices into a fresh tensor and `write_state`; `ttnn.experimental.slice_write` | |
| into a TILE state re-binds the handle and breaks traces (probe P13). | |
| **Per-stream state (PLAN.md D16).** A trace bakes one set of state addresses, so several `stream.id`s share them: | |
| `add_state(..., banks=S_MAX)` allocates `S_MAX` DRAM banks of the state before capture, `save_state(name, i)` / | |
| `load_state(name, i)` copy between the state and bank `i` (eager `ttnn.copy` on CQ0, ordered after the enqueued | |
| replays; ~80 us for a StreamPETR-size state), and `StreamBanks` is the policy every temporal bundle uses: | |
| ```python | |
| from .ttaw.trace import StreamBanks | |
| runner.add_state("prev_bev", shape=(1, 1, 22500, 256), dtype="bfloat16", pingpong=True, | |
| banks=S_MAX if S_MAX > 1 else 0) # S_MAX = 1 at the first publish | |
| streams = StreamBanks(runner, ["prev_bev"], max_streams=S_MAX, on_full="reject", max_gap_s=2.0) | |
| fresh = streams.select(stream.get("id", "default"), reset=stream.get("reset", False), | |
| timestamp_s=stream.get("timestamp_s")) # in _prepare, under the model lock | |
| params = {"use_prev_bev": 0.0 if fresh else 1.0} # first-frame RT-dev params | |
| ``` | |
| `select` returns True when the stream starts fresh (new id, `reset=True`, timestamp backwards or a gap above | |
| `max_gap_s`) after resetting its states to `init`; a new id beyond `max_streams` raises `InputError` (HTTP 400) with | |
| `on_full="reject"` or takes the least recently used stream's bank with `"evict"` (`S_MAX=1`: restart the one state). | |
| `forget(id)`, `streams` (most recent first), `active`, `describe()` (for `model.info`). | |
| **Outputs.** A variant returns either the tensors its last ops produce (allocated during the capture, alive as long | |
| as the trace) or persistent outputs: `add_output(name, shape=...)` allocates a buffer before any capture and | |
| `ctx.write_output(name, value)` copies into it inside the trace and returns it. Persistent outputs keep their address | |
| across recaptures, can be shared by several variants (shape buckets with one readback) and are never overwritten by | |
| another variant's replay; they cost one `ttnn.copy` per output. Op-produced outputs of a variant captured later can | |
| be overwritten when an earlier-captured variant replays: read them before running another variant (`read` and | |
| `__call__` return copies), and give segments read on CQ1 while CQ0 already replays the next one persistent outputs. | |
| Warm-up and release free op outputs without `force`, so an output that shares memory with something the runner | |
| does not own (a view of a weight) is never freed by the runner. | |
| | `TraceRunner` member | signature / meaning | | |
| |---|---| | |
| | constructor | `TraceRunner(device, *, num_command_queues=None, warmup_runs=1, stage_inputs=False, forbid_cache_misses=True, alloc_tracking=None, name="model")` | | |
| | `add_input` | `(name, init=None, *, shape=None, dtype="bfloat16", layout=ROW_MAJOR, memory_config=DRAM, stage_fn=None)` -> device tensor. `init`: numpy / torch / ttnn host tensor (warm-up data), else zeros | | |
| | `add_param` | `(name, value=0.0, *, shape=(1,1,1,1), dtype="float32", layout=TILE, memory_config=DRAM)` -> device tensor; broadcasts in ttnn binary ops | | |
| | `add_state` | `(name, init=None, *, shape=None, dtype="float32", layout=TILE, memory_config=DRAM, pingpong=False, banks=0)` -> buffer or (A, B) | | |
| | `add_output` | `(name, *, shape, dtype="float32", layout=TILE, memory_config=DRAM)` -> persistent output buffer (zeros) | | |
| | `add_variant` | `(name, fn, *, warmup_runs=None)`; `fn(ctx)` returns a device tensor, a `Packed`, or a list / tuple / dict of them | | |
| | `add_eager_warmup` | `(fn)`: eager device work the model runs between replays; `fn()` runs in the first `capture()`, before any trace (refused after it) | | |
| | `capture()` | warm-up + capture of all pending variants (idempotent) | | |
| | `run(variant=None, inputs=None, params=None)` | upload + non-blocking replay; returns the device outputs | | |
| | `read(variant=None, *, cq_id=0, as_torch=False)` | blocking read into preallocated host tensors -> numpy (or torch); `Packed` -> `{name: array}` | | |
| | `__call__(variant=None, inputs=None, params=None, *, as_torch=False)` | `run` + `read` | | |
| | `upload(inputs=None, params=None)` / `replay(variant=None, n=1)` | the two halves of `run` (benches, pipelining) | | |
| | `run_eager(variant=None, inputs=None, params=None)` | same function without a trace, buffers freed before returning: replay-vs-eager bit checks | | |
| | `set_params(**values)` / `write_input(name, value)` | queue param values for the next run / upload an input now (CQ0; with 2 CQs the event the next CQ1 upload waits for is re-recorded after it) | | |
| | `reset_state(name=None, value=None)` / `read_state(name)` / `state_buffer(name)` / `phase` | state control (on CQ0, ordered after enqueued replays); `value` may be a device tensor | | |
| | `save_state(name, bank)` / `load_state(name, bank)` | state <-> bank copies on CQ0 (`add_state(banks=...)`; D16) | | |
| | `outputs(variant=None, phase=None)`, `trace_ids()`, `describe()`, `timings_ms`, `phases`, `captured` | introspection (`describe()` is JSON-able: put it in `model.info`) | | |
| | `release()` | release traces and persistent tensors (idempotent; also `__exit__`) | | |
| `TraceContext` (the `ctx` of a variant): `ctx[name]` (input, param, persistent output, or a state's read buffer), | |
| `ctx.state(name)`, `ctx.write_target(name)`, `ctx.write_state(name, value)`, `ctx.write_output(name, value) -> | |
| buffer`, `ctx.variant`, `ctx.phase`, `ctx.capturing`, `ctx.device`. Registering inputs / states / variants from | |
| inside a variant function (during warm-up or capture) raises. | |
| `StreamBanks(runner, states, *, max_streams=1, on_full="reject", max_gap_s=None)`: `.select(stream_id, *, | |
| reset=False, timestamp_s=None) -> fresh`, `.forget(id)`, `.streams`, `.active`, `.describe()` (above). | |
| `pack_outputs(tensors: dict, *, dtype="float32", align=32, row_elems=None) -> Packed`: one ROW_MAJOR tensor holding | |
| every output (typecast to `dtype`; float32 is exact for bf16 and integers < 2**24), read with ONE device-to-host copy; | |
| `TraceRunner.read` unpacks it, `Packed.layout.unpack(flat)` gives `{name: array}` (views) from any array of | |
| `layout.total` elements (`PackLayout(entries, total, rows, row_elems)`, `.shape` = the device shape; `PackEntry(name, | |
| offset, numel, shape, pitch=0)` indexes the elements in row-major order; `pitch > 0`: the tensor's rows of | |
| `shape[-1]` elements are stored `pitch` elements apart, zero-padded, and `unpack` / `entry.view(flat)` return a | |
| strided view). Layout (since 0.15.0, YOLOX PORT_LOG Q9): | |
| - packed total <= `SINGLE_ROW_MAX_ELEMS` (131072 = 512 KiB of fp32): one `[1, 1, 1, total]` row, each tensor | |
| flattened and zero-padded to `align` elements -- exactly the layout and programs of 0.1.0-0.14.0; | |
| - above it, or with `row_elems=R` (a multiple of 32): `[1, 1, rows, R]`, R = `PACK_ROW_ELEMS` (1024) by default. | |
| Each tensor starts on a row boundary and is zero-padded to whole rows (a tensor of at most 64 KiB is flattened | |
| and padded by < R elements; a larger `[.., n, c]` gets zero rows appended until `n * c` fills whole rows of R, | |
| i.e. n rounded up to a multiple of `R / gcd(c, R)`; when that wastes more than 1/8 of the tensor beyond the best | |
| alternative -- few rows of an awkward width: 0.15.0-0.18.0 packed `[1, 1, 2, 20001]` fp32 as 80 MB -- it is | |
| flattened if its padded row fits 128 KiB, else its rows are zero-padded to a pitch `p` = c rounded up to a power | |
| of two dividing R, or to R, the smallest segment: one `ttnn.pad` of the row and last dims, then the reshape into | |
| rows of R; segments stay below 1.4x the tensor + one row, typical shapes keep the contiguous layout), then the | |
| segments are concatenated on the row dim. No ROW_MAJOR page above 128 KiB is reshaped, so outputs of tens of MB | |
| pack; a flat vector `[1, 1, 1, N]` wider than that is cut into `ttnn.slice` chunks; any other tensor with rows | |
| wider than 128 KiB raises `ValueError` (give it a narrower last dim). Measured (p150b, | |
| `tests/device/test_pack_outputs_device.py`, two runs, 1CQ / 2CQ; pack = replay of the packing trace, read = D2H + | |
| unpack): 64 KB single row 0.15-0.22 + 0.14-0.18 ms; 1 MB (YOLOX 14400 x 13 fp32 head level + a 240 x 240 uint32 | |
| map) 0.69-0.75 + 0.30-0.54 ms; 16 MB (`[1, 1, 262144, 16]` fp32) 1.24-1.74 + 3.7-8.1 ms (2-4 GB/s); bit-exact | |
| through trace replays (`logs/ttaw/q9_pack_outputs_device*.log`). | |
| Why: a single ROW_MAJOR row is ONE page, and `ttnn.reshape` / `concat` stage whole pages in L1 (2 x the | |
| destination page per kernel copy for pages that are not 16-byte aligned), so the 0.14.0 layout failed above | |
| ~0.6 MB with `TT_FATAL: RM reshape dest staging does not fit in L1`. | |
| `CQ_COMPUTE = 0`, `CQ_INPUT = 1`. | |
| **Alloc tracking (debug).** `TT_METAL_TRACE_ALLOC_TRACKING=1` must be exported before Python imports ttnn; then | |
| `ttnn.execute_trace` refuses to replay while a buffer allocated after a capture is alive. `alloc_tracking_enabled()` | |
| reports it; `TraceRunner(alloc_tracking=True)` raises if it is off. The runner acknowledges the outputs of every trace | |
| captured after the first (plain ttnn would flag them: `tests/device/test_trace_alloc_tracking_device.py`). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 4. Tensors (C03, `ttaw.tensors`) | |
| | name | meaning | | |
| |---|---| | |
| | `TILE`, `round_up(n, multiple=32)`, `tile_padded_shape(shape)`, `pad_to_tile(a, value=0)` | tile geometry (host) | | |
| | `pad_to_capacity(a, capacity, *, axis=0, value=0, truncate=False) -> (padded, n_valid)` | fixed-capacity buffers (capacities are grid-independent constants) | | |
| | `as_4d(a)` | prepend unit dims to rank 4 | | |
| | `float32_to_bf16_bits`, `bf16_bits_to_float32`, `round_to_bf16` | RNE bf16 conversion identical to torch (expected values for exact tests) | | |
| | `ttnn_dtype(name)`, `dtype_name(dtype)` | `"bf16"`, `"fp32"`, `"bfp8"`, `"bfp4"`, `"uint32"`, `"int32"`, `"uint16"`, `"uint8"` aliases | | |
| | `to_host_tensor(value, dtype, layout=None, *, shape=None)` | numpy / torch / ttnn host tensor -> ttnn host tensor; integers never pass through a float intermediate (int64 input to `ttnn.from_torch` would) | | |
| | `to_device(value, device, dtype="bf16", layout=None (TILE), memory_config=None (DRAM))` | eager upload | | |
| | `to_numpy(t)` | ttnn (device or host) / torch -> numpy (bf16 -> float32) | | |
| | `to_layout(t, layout)`, `typecast(t, dtype)`, `to_fp32(t)`, `to_bf16(t)` | no-ops when nothing changes (no extra program) | | |
| | `fp32_island(fn, *tensors, out_dtype="bfloat16")` | run `fn` on fp32 copies, cast results back | | |
| | `upload_u8(array, device, *, out_dtype="bfloat16", layout=None, memory_config=None)` | uint8 upload + on-device typecast (exact); in a trace keep the uint8 tensor as the input and `ttnn.typecast` as the first op | | |
| | `HostStaging(shape, dtype)` | persistent ROW_MAJOR host tensor; `.write(array)` copies into its buffer through a `torch.from_dlpack` alias (`zero_copy=True` on this tree for float32 / bfloat16 / uint32 / int32 / uint16 / uint8) and returns `.tensor`; pass it as a `run(inputs=...)` value | | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 5. Weights (C04, `ttaw.weights`) | |
| ```python | |
| from .ttaw.weights import OnnxWeights, WeightCache, fold_bn_conv | |
| w = OnnxWeights(weights_path / "pts_backbone_neck_head_centerpoint.onnx") # parsed as data, never executed | |
| for node in w.nodes("Conv", "/backbone/*"): # graph order, glob or regex | |
| wf, bf = w.fold_conv_bn(node.name) # + its BatchNormalization consumer, fp64, rounded once | |
| k = w.param("/dit/blocks.0/attn/MatMul", 1) # anonymous initializer, addressed by consuming node + slot | |
| cache = WeightCache("tt_centerpoint/base", w.sha256, version=f"{__version__}.prep1") # $TT_CACHE_PATH | ~/.cache/ttaw | |
| tw = cache.get("backbone.0.conv.w", lambda: wf, dtype="bfloat16", layout=ttnn.ROW_MAJOR_LAYOUT, device=dev) | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `OnnxWeights(path, *, load_external_data=True)` | `.array(name)` (initializer / Constant, through Identity), `.has`, `.initializer_names()`, `.state_dict()`, `.node(name)`, `.nodes(op_type=None, pattern=None, *, regex=False)`, `.find_node(pattern, op_type=None, *, regex=False)`, `.producer(t)`, `.consumers(t)`, `.consumer_of(t, op_type=None)`, `.param(node, slot)`, `.params(node)`, `.conv(node) -> ConvParams`, `.gemm(node) -> GemmParams`, `.matmul_weight(node)`, `.batchnorm(node) -> BatchNorm`, `.fold_conv_bn(conv, bn=None, *, dtype=np.float32)`, `.sha256`, `.input_names`, `.output_names`, `.opset`. External data outside the model directory is refused. Unnamed nodes are `"<OpType>_<index>"` | | |
| | `OnnxNode`, `ConvParams` (weight, bias, strides, pads, dilations, group, kernel_shape, output_padding, auto_pad, output_shape), `GemmParams` (weight as stored, bias, trans_a, trans_b, alpha, beta) | frozen records; `pads` is the explicit attribute: with `auto_pad` `SAME_*` derive the padding from the input size | | |
| | `BatchNorm(gamma, beta, mean, var, eps=1e-5)` | float64; `.scale`, `.shift`, `.channels`, `.apply(x, axis=1)`, `.from_state_dict(sd, prefix, eps)` | | |
| | `fold_bn(w, b, bn, *, axis=0, dtype=np.float32)` | generic fold (`dtype=None` keeps float64 for a single final rounding to the device dtype) | | |
| | `fold_bn_conv(w, b, bn)` | Conv weight `[Cout, Cin/g, k...]`, any groups | | |
| | `fold_bn_linear(w, b, bn, *, layout="out_in")` | torch Linear / Gemm transB=1 (`out_in`) or MatMul / Gemm transB=0 (`in_out`); also BN1d | | |
| | `fold_bn_conv_transpose(w, b, bn, *, groups=1)` | ConvTranspose weight `[Cin, Cout/g, k...]` | | |
| | `load_safetensors(path)`, `load_torch_checkpoint(path, *, key=None)` (`torch.load(weights_only=True)`), `load_state_dict(path)` -> `StateDict` | `.safetensors`, `.pth/.pt/.ckpt/.bin`, `.npz` | | |
| | `StateDict(tensors)` | `.sub(prefix)`, `.strip(prefix)`, `.bn(name, eps)` | | |
| | `WeightCache(namespace, source_digest="", *, version="", root=None, enabled=None)` | `.get(key, make, *, dtype, layout=None, device=None, memory_config=None, shape=None)` caches ttnn host tensors as `.tensorbin` (atomic writes; rebuilt if unreadable or of another dtype / layout / `shape`), `.path(...)`, `.clear()`, `.hits`, `.misses`; `TTAW_WEIGHT_CACHE=0` disables. The cache outlives images (`/tensor-cache` is the host's `~/.cache/tt-model/<name>/tensors`): pass `version` (bundle `__version__` + a prep revision) and bump it whenever the code that prepares the tensors changes | | |
| | `cache_root()`, `file_sha256(path)` | helpers | | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 6. Precision and knobs (C05, `ttaw.precision`, `ttaw.knobs`) | |
| `ttnn.matmul` / `linear` drop to LoFi when a `program_config` or `core_grid` comes without a compute config, and | |
| `ttnn.WormholeComputeKernelConfig()` without `math_fidelity` is `MathFidelity.Invalid`. Always pass one: | |
| ```python | |
| from .ttaw.precision import PrecisionPolicy, compute_kernel_config | |
| POLICY = PrecisionPolicy({"head.reg*": "accurate", "backbone.*": "HiFi2+fp32:w=bfp8"}, default="balanced") | |
| policy = POLICY.with_env("CENTERPOINT") # CENTERPOINT_PRECISION="backbone.*=HiFi4+fp32;*=HiFi2+fp32" | |
| y = ttnn.linear(x, w, compute_kernel_config=policy.compute_kernel_config("head.reg.fc1"), program_config=pc) | |
| w_dtype = policy.resolve("backbone.block3.conv2").weights_dtype() | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `compute_kernel_config(fidelity="HiFi2", *, fp32_acc=True, approx=False, packer_l1_acc=False, dst_full_sync=False)` | explicit `ttnn.WormholeComputeKernelConfig` (= `BlackholeComputeKernelConfig`) | | |
| | `Precision(fidelity="HiFi2", fp32_acc=True, approx=False, packer_l1_acc=False, dst_full_sync=False, weights="bfloat16", activations="bfloat16")` | `.parse("HiFi4+fp32+approx+l1acc+fullsync:w=bfp8:a=bf16" or preset)`, `.label`, `.with_(**)`, `.compute_kernel_config()`, `.weights_dtype()`, `.activations_dtype()` | | |
| | `PRESETS` | `accurate` (HiFi4 + fp32), `balanced` (HiFi2 + fp32, the default), `fast` (LoFi, bfp8 weights; only after gates pass) | | |
| | `PrecisionPolicy(rules, default="balanced")` | first matching glob wins; `.resolve(module)`, `.compute_kernel_config(module)` (cached), `.override(spec)`, `.with_env(prefix, env=None)`, `.describe()` (rules + which modules resolved to what) | | |
| | `FIDELITIES` | `("LoFi", "HiFi2", "HiFi3", "HiFi4")` | | |
| Other silent defaults to override by hand: `ttnn.layer_norm` epsilon 1e-12, SDPA `is_causal=True`, `ttnn.embedding` | |
| PADDED returning the cached pad row, HARDSWISH fused into conv2d skipped. | |
| Knobs (one per optimization, default = measured best, pinned in `tt-model.yaml serve.env`): | |
| ```python | |
| from .ttaw.knobs import Knob, Knobs | |
| KNOBS = Knobs("CENTERPOINT", [Knob("FUSED_HEAD", True, "merged head convs"), | |
| Knob("ACT_BLOCK_H", 64, "conv act_block_h", choices=(32, 64, 128)), | |
| Knob("BFP8_WEIGHTS", False, "bfp8 backbone weights", experiment=True)]) | |
| knobs = KNOBS.read() # once, in _build; env CENTERPOINT_FUSED_HEAD=0 is the A/B switch | |
| assert KNOBS.serve_env() == yaml_serve_env_subset # host test: the image pins the defaults | |
| ``` | |
| `Knob(name, default, doc="", choices=None, experiment=False)`; `Knobs(prefix, knobs)`: `.read(env=None, | |
| **overrides) -> KnobValues` (attribute / item access, `.source(name)`, `.overridden()`, `.as_dict()`, immutable; | |
| warns when an experiment knob is set), `.defaults()`, `.serve_env(values=None)`, `.doc_table()`, `.env_name(name)`; | |
| `parse_bool(text)`. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 7. Metrics, goldens, gates (C06, `ttaw.metrics`, `ttaw.golden`) | |
| ```python | |
| # tests/test_pcc_device.py of a bundle | |
| from ..ttaw.golden import GateRegistry, compare_taps, load_goldens | |
| from ..ttaw.metrics import pcc, match_detections | |
| GATES = GateRegistry.for_test(__file__, {"backbone": 0.999, "head.heatmap": 0.99, | |
| "dets.recall": (0.95, "min", "recall"), "plan.ade": (0.5, "max", "ade")}) | |
| def test_taps(device_taps): # numpy dict from an eager device run or trace outputs | |
| with load_goldens(SPEC_DIR / "golden/sample0.npz") as gold: | |
| report = compare_taps(device_taps, gold, gates=GATES, names=["backbone", "head.heatmap"]) | |
| report.save_json(LOGS / "compare_backbone.json"); print(report.table()) | |
| assert report.passed | |
| ``` | |
| `GateRegistry` stores frozen gates in `<test stem>.gates.json` next to the test at the first green run. A later | |
| declaration that is looser (lower `min`, higher `max`, other direction or metric) raises `GateLoosenedError`; | |
| tighter ones are re-frozen; failing checks freeze nothing. `TTAW_GATES_READONLY=1` (or `write=False`) never writes. | |
| Changing a frozen gate = editing the JSON by hand + disclosure in VERIFICATION_<date>.md. | |
| | name | meaning | | |
| |---|---| | |
| | `pcc(t, r)` | float64 Pearson; NaN on non-finite input; a constant side (detected exactly, max == min) gives 1.0 for two equal constants (`np.allclose`) and 0.0 otherwise, also for two different constants | | |
| | `masked_pcc(t, r, mask)`, `valid_row_pcc(t, r, n_valid=None, *, rows=None, axis=0)` | padded / masked buffers | | |
| | `error_stats(t, r)` | `{pcc, max_abs, mean_abs, rel_l2, n}` | | |
| | `argmax_agreement(t, r, *, axis=-1, mask=None)`, `label_agreement(t, r, *, mask=None)` | class decisions (ties -> lowest index) | | |
| | `mask_iou(a, b)`, `mean_iou(t, r, num_classes, *, ignore_index=None) -> (miou, per_class)`, `box_iou_xyxy(a, b)` | IoU | | |
| | `topk_overlap(t_idx, r_idx, k=None)`, `topk_set_overlap(t_scores, r_scores, k)` | data-dependent selections | | |
| | `match_detections(t_centers, t_labels, t_scores, r_centers, r_labels, r_scores, *, max_dist=0.5, dims=2, same_label=True) -> DetectionMatch` | greedy same-label BEV-centre matching (`.pairs`, `.recall`, `.precision`, `.matched`, `.center_errors`, `.score_errors`, `.to_dict()`); box arrays `[N, 7]` work as centres; a box with a non-finite centre never matches | | |
| | `ade_fde(pred, ref, *, dims=2) -> (ade, fde)` | trajectories `[..., T, D]` | | |
| | `as_array(x)` | numpy / torch (bf16 ok) / list -> numpy | | |
| | `save_goldens(path, tensors, meta=None, *, compress=False)`, `load_goldens(path) -> Goldens` | `.npz` + JSON `__meta__`; `Goldens` is a lazy mapping with `.meta`, `.sha256`, `.close()`, context manager | | |
| | `TapRegistry(enabled=True, *, include=("*",), exclude=())` | `.tap(name, value) -> value` (device tensors read at once; raises inside a capture), `.scope(prefix)`, `.wants`, `.names()`, `.to_dict()`, `.save(path, meta)`, `.clear()`; `NULL_TAPS` is a disabled registry | | |
| | `Gate(threshold, direction="min", metric="pcc")` | `.of(spec, metric=None)` (a bare number takes the direction of `metric` from `DIRECTIONS` and is refused for a metric of unknown direction; a direction contradicting a known metric, e.g. a `"max"` PCC gate, is refused), `.passes(v)` (NaN fails), `.looser_than(other)` | | |
| | `GateRegistry(path, gates=None, *, write=None)` | `.for_test(test_file, gates)`, `.check(name, value, gate=None) -> GateResult`, `.require(...)` (AssertionError), `.gate(name)`, `.names()`, `.frozen`, `.results`, `.save()` | | |
| | `CompareReport(title="", meta=None, gates=None)` | `.add(name, test, ref, *, metric="pcc", gate=None)`, `.add_value(name, value, *, metric, gate=None, stats=None)`, `.passed`, `.table()`, `.to_dict()`, `.save_json(path)`; `TapComparison` records | | |
| | `compare_taps(test, golden, *, gates=None, metric="pcc", names=None, title="", meta=None)` | one report over common (or named) taps; missing taps fail when gated | | |
| | `METRICS` | name -> (function, direction) used by reports (`pcc`, `masked_pcc`, `argmax_agreement`, `label_agreement`, `mask_iou`, `max_abs`, `mean_abs`, `rel_l2`) | | |
| | `DIRECTIONS` | metric name -> `"min"` / `"max"` for every gate metric that may be given as a bare number: the `METRICS` plus `recall`, `precision`, `iou`, `miou`, `agreement`, `topk_overlap` (min) and `ade`, `fde`, `center_err*`, `score_err*`, `max_abs_err` (max) | | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 8. Bench and profile (C07, `ttaw.profiling`) | |
| ```python | |
| from .ttaw.profiling import AiclkSampler, bench_trace_runner, signposted, summarize_ops | |
| with AiclkSampler(interval_s=0.05) as clk: # sysfs tt_aiclk + hwmon power / temperature | |
| bench = bench_trace_runner(model.runner, "default", {"pillars": p}, iters=100, post=model.decode) | |
| bench.save_json(f"{ROOT}/logs/centerpoint/bench_baseline.json", config=model.device_info, aiclk=clk.summary()) | |
| print(bench.table()) # host_in / h2d / trace / d2h / post / e2e / b2b, p50 p99 | |
| with signposted("trace"): # profile_ops.py under `python -m tracy -r -p -v -o ...` | |
| model.runner.replay("default"); ttnn.synchronize_device(dev) | |
| print(summarize_ops(f"{ROOT}/generated/profiler/centerpoint_baseline").table()) # ops, kernel sum, op2op, span | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `StageBench(name="")` | `.stage(name)` context, `.add(name, ms)`, `.summary()` (`{stage: {n, p50, p99, mean, min, max}}`), `.table()`, `.to_dict(**extra)`, `.save_json(path, **extra)`; `STAGES` order | | |
| | `bench_trace_runner(runner, variant, inputs, *, params=None, iters=100, warmup=10, post=None, b2b_iters=None, name="")` | the standard stage breakdown of a `TraceRunner` variant | | |
| | `time_b2b(enqueue, sync, n=100, warmup=5) -> ms` | back-to-back device time per iteration | | |
| | `signpost(name, message=None) -> bool`, `signposted(name)` | Tracy markers (no-op without `tracy`) | | |
| | `read_device_profiler(device)` | `ttnn.ReadDeviceProfiler` per trace segment (the buffer holds ~1000 ops) | | |
| | `summarize_ops(source, *, start="trace", end=None, last_replay_session=None, freq_mhz=None, top_gaps=10) -> OpsSummary` | `ops_perf_results*.csv` / `cpp_device_perf_report.csv` / directory / rows; `OpsSummary`: `ops`, `kernel_sum_us`, `fw_sum_us`, `op2op_sum_us`, `span_us`, `by_op`, `fidelity`, `top_gaps`, `.table(top)`, `.to_dict()`. CLI: `python -m <pkg>.ttaw.profiling <csv|dir> [--start trace] [--end trace_end] [--json out]` | | |
| | `find_ops_csv(path)` | newest ops CSV under a directory | | |
| | `AiclkSampler(chip=0, interval_s=0.05, *, root="/sys/class/tenstorrent")` | `.start()`, `.stop()`, context manager, `.sample_once()`, `.summary()`, `.available` | | |
| `common/tools/p2_trace_bench.py --mode eth-1cq|eth-2cq|eth-2cq-staged|worker-2cq` measures the runner itself. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 9. I/O and outputs (C08, `ttaw.io`, `ttaw.outputs`) | |
| `ttaw.io` (numpy only; every client mistake raises `InputError`, which the server maps to HTTP 400): | |
| `b64decode(s, *, field, max_bytes)`, `PointCloud(points, fields, frame_id)` (`.select(names, fill=None)`; rows with a | |
| non-finite `x` / `y` / `z` column -- the first three columns when unnamed -- dropped), | |
| `decode_points(spec, *, max_points=2_000_000, max_bytes=None, default_fields=DEFAULT_POINT_FIELDS)`, | |
| `load_points(source, *, fields=None, fmt=None, frame_id="base_link", default_fields=...)` (path `.bin/.npy/.npz/.pcd`, | |
| bytes + `fmt`, arrays, structured arrays, envelopes), `decode_image(raw, *, fmt="auto", max_side=8192)`, | |
| `load_image(image)` -> RGB uint8 HxWx3, `CameraImage`, `decode_cameras(images, calibration=None, *, order=None, | |
| require_calibration=True, max_bytes=None)` (the server's `images[]`), its Python-API twins (0.15.0) | |
| `load_camera(source, *, name=None, calibration=None, require_calibration=False)` (a `CameraImage`, `{"camera", | |
| "image" | "path" | "data", "intrinsics", "T_ref_from_camera", "distortion", "timestamp_s"}` or any `load_image` source | |
| with `name`; inline calibration wins over `calibration["cameras"][name]`) and `load_cameras(images, calibration=None, | |
| *, order=None, require_calibration=True, calib_dir=None)` (a list or a `{name: image}` mapping; `{"preset": name}` | |
| resolved in `calib_dir`; reordered to `order`, every camera exactly once: `model(images=..., calibration=...)` and | |
| `/predict` see the same cameras), `decode_rois(rois, *, cameras=None, labels=None, max_rois=4096)` (2-D detections | |
| as an input, `[{"camera", "label", "score", "box_xyxy"}]`: names / indices checked against `labels`, score in [0, 1], | |
| finite ordered boxes; PointPainting `rois`, BUNDLE_CONVENTIONS.md 7.2), `parse_transform(obj, *, field)` (4x4 / 3x4 / `{matrix}` / | |
| `{translation, rotation_wxyz|rotation_xyzw|rotation}` / `{x, y, z, roll, pitch, yaw}` tf2 RPY), | |
| `parse_intrinsics(obj, *, field)`, `resolve_calibration(calibration, calib_dir)` (`{"preset": name}` -> | |
| `calib_dir/<name>.json` merged), `decode_named_arrays(spec, schema=None, *, max_bytes=None)` (planner inputs), | |
| `check_named_arrays(arrays, schema)` (names, shapes, numeric and finite values, integral values in range for integer | |
| dtypes; every problem is an `InputError`), `load_named_arrays(source, schema=None)` (mapping / envelope / `.npz` | |
| path or bytes: the Python-API form), | |
| `encode_array(arr, *, fmt="npz"|"npz_compressed"|"raw"|"list", key)`, `encode_png(image, *, key)`, `to_jsonable(obj)` | |
| (float32 / float16 scalars rounded to 6 significant digits, float64 exact), | |
| `POINT_FORMATS`, `DEFAULT_POINT_FIELDS = ("x", "y", "z", "intensity")`, `MAX_POINTS_DEFAULT`. | |
| `ttaw.outputs` (each has `to_dicts()` and `to_dict(output_format="json"|"npz")` = the `/predict` body with `model`, | |
| `frame_id`, `meta`, `timing_ms`): | |
| | class | fields | | |
| |---|---| | |
| | `Detections3D(boxes [N,7] x y z l w h yaw, scores, label_ids, velocities=None, labels=(), model="", frame_id="base_link", timing_ms, meta)` | rows sorted by score at construction; `detections[] {label, label_id, score, center, size, yaw, velocity}`; npz adds `arrays` | | |
| | `Detections2D(boxes_xyxy, scores, label_ids, labels=(), extras={}, model, frame_id="camera", ...)` | `detections[] {label, label_id, score, box_xyxy}`; `extras` (e.g. `{"semseg": Mask2D}`) encoded by name | | |
| | `Segmentation3D(label_ids [N], class_names=(), scores=None, ...)` | `labels` (npz uint8 / uint16), `class_names`, `class_counts`, `scores` | | |
| | `Mask2D(mask (H, W), class_names=(), encoding="png", ...)` | `mask` (png, or npz for `output_format="npz"` / non-uint8), `class_counts`; `.payload(fmt)` | | |
| | `Trajectory(poses [T, D], columns=("x", "y", "yaw"), turn_indicator=None, predicted_agents=None, ...)` | `trajectory`, `columns`, `num_poses`, `turn_indicator`, `predicted_agents` (npz) | | |
| `label_name(labels, id)` maps ids to names (the id as text when out of range). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 10. `ModelBase` (C08, `ttaw.api_base`) | |
| ```python | |
| from .ttaw.api_base import ModelBase | |
| from .ttaw.outputs import Detections3D | |
| class CenterPoint(ModelBase): | |
| MODEL_NAME = "centerpoint-p150"; ENV_PREFIX = "CENTERPOINT" | |
| DEFAULT_REPO = "AutowareFoundation/lidar_centerpoint"; DEFAULT_TAG = "v4.1" | |
| DEFAULT_REVISION = "494c8171def40bd36cc2feb323e0a5acbfab132b"; ALLOW_PATTERNS = ["base/*", "tiny/*"] | |
| VARIANTS = ("base", "tiny"); DEFAULT_VARIANT = "base"; INPUT_KIND = "lidar" | |
| LABELS = ("CAR", "TRUCK", "BUS", "BICYCLE", "PEDESTRIAN") | |
| RUNTIME_PARAMS = {"score_threshold": (float, 0.0, 1.0, 0.35), "max_detections": (int, 1, 1000, 500)} | |
| DEVICE_DEFAULTS = {"num_command_queues": 1, "trace_region_size": 64 << 20} | |
| def _build(self): # weights -> device tensors, TraceRunner + inputs / params / states / variants | |
| self.runner = TraceRunner(self.device, name=self.MODEL_NAME); ... | |
| def _warm_one(self, v): # capture (default_warmup_variants() -> [{"variant": self.variant}]) | |
| self.runner.capture() | |
| def _prepare(self, points=None, **kw): # host pre-processing; raise io.InputError for client mistakes | |
| ... | |
| def _forward(self, prep): # under the model lock | |
| return self.runner("default", inputs={...}) | |
| def _postprocess(self, raw, prep, params) -> Detections3D: | |
| return Detections3D(boxes, scores, ids, labels=self.LABELS, model=self.MODEL_NAME) | |
| def _release(self): | |
| self.runner.release() | |
| def extra_info(self): | |
| return {"trace": self.runner.describe(), "knobs": self.knobs.as_dict()} | |
| with CenterPoint.from_pretrained() as model: # weights (pinned) -> device (ETH, 12x10) -> build -> warm | |
| out = model("samples/test.npz", score_threshold=0.4) | |
| print(out.to_dict()["num_detections"], out.timing_ms, model.info) | |
| ``` | |
| `from_pretrained(model_id=None, *, revision=None, variant=None, device_id=None, device=None, dispatch=None, | |
| num_command_queues=None, weights_dir=None, warmup_variants="default", verbose=False, **compile_params)`: weights are | |
| resolved first (`weights_dir` > `<ENV>_WEIGHTS_DIR` > local dir `model_id` > HF snapshot at the pinned revision, | |
| offline fallback to the cache), then the device opens with `DEVICE_DEFAULTS` < `<ENV>_*` / `TT_DEVICE_ID` < explicit | |
| arguments; a caller-provided `device` is left open by `close()`. Variant default: `<ENV>_VARIANT` or | |
| `DEFAULT_VARIANT`. A failing build closes everything. | |
| Other members: `resolve_weights(...)` (also module-level `resolve_weights(model_id, *, revision, allow_patterns, | |
| weights_dir, env_var)`), `device_config(**overrides)`, `requires_calibration()`, `validate_params(params)` (unknown / | |
| outside `[min, max]` (either bound may be None) / wrong type -> `InputError`; bools must be real booleans, ints must | |
| be integral and not booleans, floats finite; numeric strings are accepted), `warmup(variants="default")` (idempotent), | |
| `default_warmup_variants()`, `__call__(points=None, *, sweeps, images, calibration, stream, inputs, **params)` / | |
| `predict` (one re-entrant lock around prepare + forward + postprocess; fills `timing_ms` preprocess / device / | |
| postprocess / total), `info`, `closed`, `close()` (idempotent, also at interpreter exit), context manager. | |
| Class attributes also: `WEIGHTS_LICENSE`, `CAMERA_ORDER`, `REQUIRE_CALIBRATION` (None: true for multicam), | |
| `POINT_FIELDS`, `EXTRA_INPUTS` (model-specific input keywords handed to `_prepare` instead of being validated as | |
| runtime params, e.g. PointPainting `("rois",)`; the server's `decode_extra` adds them to the call kwargs) and | |
| `INPUT_SCHEMA` (`{name: (shape with None for free dims, dtype)}`: `inputs=` is decoded with | |
| `io.load_named_arrays` and checked on every call, API and server alike). `RUNTIME_PARAMS` may not reuse an input | |
| keyword (`from_pretrained` raises `TypeError`). `info` adds `extra_inputs` and `input_schema`; `__call__` re-checks | |
| `closed` once it holds the lock. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 11. Server (C08, `ttaw.server`) | |
| `create_app(spec: ServerSpec, *, model_factory=None) -> FastAPI` implements BUNDLE_CONVENTIONS.md section 7: | |
| `GET /` (routes), `GET /health` and `/v1/health` (always 200: `ok` / `starting` / `error`, plus device), | |
| `GET /info` (model, task, io, autoware, weights, device, input limits, calibration presets, labels, variants, | |
| runtime / compile params, warm-up and boot ms, source), `GET /v1/models` (stub), `POST /predict`. Errors: 400 | |
| (`InputError`, bad params, body above `<ENV>_MAX_BODY_MB` by Content-Length), 422 (schema: unknown fields are | |
| refused), 503 (starting / failed boot), 500 (`inference failed: <Type>: <msg>`). The lifespan reads the environment | |
| once (`config_from_env`), refuses a `TT_MESH_SHAPE` other than 1x1, builds the model through | |
| `spec.model_cls.from_pretrained` (or `model_factory(cfg)`), and closes it on shutdown under the lock. A body holding a | |
| NaN / infinity (strict JSON has none) is a 500 `inference failed: non-finite value at body.<path>`. `/info` `input` | |
| also lists `extra_inputs` and `input_schema`. | |
| `ServerSpec(model_name, env_prefix, model_cls=None, task="", default_weights="", owner="changh95", hardware=HARDWARE, | |
| io="", autoware={}, source={}, calib_dir=None, default_variant=None, version="0.1.0", description="", | |
| request_model=PredictRequest, decode_extra=None, info_extra=None)`. Model-specific request fields: subclass | |
| `PredictRequest` (e.g. `rois: Optional[List[...]] = None`), add them to the call kwargs in | |
| `decode_extra(request, kwargs, max_bytes)` and list them in the model's `EXTRA_INPUTS`. `app.state.ttaw` | |
| (`ServerState`) holds `model`, `config`, `ready`, `error`, `lock`, `model_factory` (tests: assign a stub before | |
| `TestClient(app)`) and `predict` (the route handler, for API == server checks). Also exported: | |
| `parse_mesh_shape`, `config_from_env(spec, env=None)`, `default_model_factory(spec)`, `decode_request(req, spec, | |
| max_bytes)`, the request models `PointsSpec`, `SweepSpec`, `ImageSpec`, `StreamSpec`, `ArraysSpec`, | |
| `PredictRequest`. | |
| `ttaw.server.client` and `ttaw.server.smoke` are standard-library only and import nothing from ttaw, so a bundle's | |
| `smoke_test.py` (run by `container_smoke.sh` with the host's `python3`, which has no numpy) loads them by path: | |
| ```python | |
| HERE = Path(__file__).resolve().parent # code/<pkg>/server | |
| sys.path.insert(0, str(HERE.parent / "ttaw" / "server")) | |
| import client as ttaw_client | |
| import smoke as ttaw_smoke | |
| health = ttaw_client.wait_ready(url, wait_s=600) | |
| info, fails = ttaw_client.check_service(url, expect_dispatch="eth", expect_grid="12x10") # PLAN 0.3 item 6 | |
| pinned = ttaw_smoke.pinned_config(f"{staged}/tt_kernel_manifest.json", profile) # serve.env + profile | |
| fails += ttaw_smoke.check_pinned(info, pinned, "CENTERPOINT") # /info runs the pins | |
| code, body = ttaw_client.post(url, ttaw_client.build_request(points=str(SAMPLE))) | |
| if ref := ttaw_smoke.find_reference(SAMPLE, pinned["profile"], info.get("variant")): # stored CPU output | |
| metrics, more = ttaw_smoke.compare_with_reference(body, ttaw_smoke.load_json(ref), gates={"min_recall": 0.95}) | |
| fails += more | |
| fails += [f] if (f := ttaw_client.check_bad_request(url)) else [] | |
| print(ttaw_client.smoke_line("centerpoint-p150", f"n={body.get('num_detections')}", fails)) | |
| ``` | |
| Functions: `build_request(points=None, fields=None, images=(), calib=None, calib_preset=None, inputs=None, | |
| params=None, stream_id=None, reset=False, output_format="json")`, `post(url, payload, timeout=120)`, | |
| `get_json(url, timeout=10)`, `wait_ready(base, wait_s=0, poll_s=5)`, `check_service(base, *, expect_dispatch="eth", | |
| expect_grid="12x10") -> (info, failures)`, `check_bad_request(base, payload=None)`, `smoke_line(model, summary, | |
| failures)`, `b64file`, `main` (CLI: `--points/--image/--calib/--calib-preset/--inputs/--param/--out/--url`). | |
| `ttaw.server.smoke` functions: | |
| | name | meaning | | |
| |---|---| | |
| | `compare_with_reference(body, reference, gates=None) -> (metrics, failures)` | served `/predict` body vs the stored CPU-reference body of the same input, by the reference's keys: `detections` (greedy same-label matching in descending reference score; 3-D rows by BEV centre distance, 2-D rows by IoU; recall, precision, max score difference), `trajectory` (ADE / FDE over x, y), every encoded array (integer: fraction of equal elements; float: max abs difference, equal infinities agree, a NaN fails), `model` / `frame_id` equal. Undecodable input is a failure, never an exception | | |
| | `DEFAULT_GATES` | `max_center_dist` 0.5 m, `min_iou` 0.5, `min_recall` 0.95, `min_precision` 0.95, `max_score_err` 0.05, `min_label_agreement` 0.99, `max_abs_err` None (report only), `max_ade` 0.5, `max_fde` 1.0; `gates=` overrides by name (unknown names raise) | | |
| | `find_reference(sample, *keys)` | first existing `<stem>.<key>.reference.json` (keys: serve profile, model variant), else `<stem>.reference.json`, else None | | |
| | `pinned_config(manifest, profile=None)` | `{"profile", "env", "weights"}` of a staged `tt_kernel_manifest.json`: `serve.env` with the profile's env on top (tt-model's merge), default profile when None | | |
| | `check_pinned(info, pinned, prefix)` | `/info` agrees with the pinned `<prefix>_DISPATCH`, `_NUM_CQS`, `_VARIANT` and the weights revision (absent pins are not checked) | | |
| | `decode_array(spec) -> Array(shape, dtype, values)` | numpy-free decoding of `io.encode_array` (npz, npz_compressed, raw, list; C or Fortran order, either byte order) and `io.encode_png` (8-bit grey / RGB(A), every PNG filter) | | |
| | `describe_metrics(metrics)`, `load_json(path)`, `ARRAY_FORMATS` | helpers | | |
| The reference body is the `to_dict()` of the fp32 CPU reference on the shipped sample, stored next to it as | |
| `<stem>.reference.json` (per profile or variant when the output depends on it). Synthetic samples may give no | |
| detections, so agreement with that reference, not a detection count, is their smoke gate (PLAN.md section 6.3). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 12. Vendoring (C09, `common/tools/vendor.py`) | |
| `python common/tools/vendor.py <bundle> [--pkg tt_<model>] [--rev REV] [--allow-dirty] [--check] [--dry-run]`. | |
| - **What is vendored is a committed revision**: `ttaw/` as committed at `--rev` (default `HEAD`), read with `git | |
| archive`, never the working tree of `common/` (other agents' uncommitted modules live there: YOLOX PORT_LOG Q7). | |
| Uncommitted changes under `common/ttaw` (staged, unstaged, untracked) are listed in a WARNING and left out: | |
| commit your change (`git add <your paths>`, bump the version + CHANGELOG), then vendor. The revision is resolved | |
| once, so a commit landing meanwhile never mixes into the recorded `source_commit`. | |
| - `--allow-dirty` vendors the working tree instead, for a local experiment only: `source_dirty` is true when it had | |
| uncommitted changes or its files differ from `ttaw/` of HEAD (.gitignore'd files included: `source_dirty: false` | |
| always means "verified against `source_commit`"); `check_bundle.py` warns about such a copy and refuses it at | |
| `--stage publish`. Outside a git checkout the working tree is vendored with `source_commit: null` (same | |
| treatment). | |
| - The destination must be a plain directory: one that is, or reaches through a symlink, `common/ttaw` (or an | |
| ancestor of it) is refused before anything is written (mirroring HEAD there would delete other agents' untracked | |
| files and revert their uncommitted changes). | |
| - The copy mirrors that tree into `code/<pkg>/ttaw` (stale files removed; `__pycache__`, `*.pyc`, `*.bin`, `*.log` | |
| never copied) and writes `VENDORED.json`: `version` (of the vendored `__init__.py`), `source_commit` (full sha), | |
| `source_ref` (the `--rev` given), `source_tree` (`commit` | `working-tree`), `source_dirty`, `vendored_at`, | |
| per-file sha256. `research/packaging/scripts/instantiate_bundle.py --out` vendors the same way (`--vendor-rev`, | |
| `--vendor-allow-dirty`; an existing copy is kept unless `--force`, and always with `--only <paths>`). | |
| - `--check` (exit 1 on any line): "modified in bundle" / "missing in bundle" / "extra in bundle" (the copy vs its | |
| `VENDORED.json`); "not at recorded revision <sha>" (`VENDORED.json` is not `ttaw/` of `source_commit`; | |
| "unverified:" when that commit is not in `common/`; skipped for a `source_dirty` copy); "outdated:" (`VENDORED.json` | |
| vs `ttaw/` committed at `--rev`, default HEAD: re-vendor and re-run the gates). Uncommitted work in `common/ttaw` | |
| never makes a bundle look outdated. `check_bundle.py` reports "outdated:" and "unverified:" as warnings (errors at | |
| `--stage publish`), every other line as an error. | |
| Never edit the vendored copy: change `common/ttaw`, bump `__version__` + CHANGELOG, commit, re-vendor every consumer | |
| and re-run their gates (PLAN.md section 5.1 step 5). Functions: `vendor(bundle, pkg=None, *, dry_run=False, rev=None, | |
| allow_dirty=False)`, `check(bundle, pkg=None, *, rev=None)`, `load_source(rev=None, *, allow_dirty=False) -> Source`, | |
| `rev_files(rev)`, `rev_hashes(rev)`, `rev_version(rev)`, `resolve_rev(rev)`, `dirty_paths()`, `tree_hashes(root)`, | |
| `source_version()` (working tree), `git_state()`, `find_package(bundle, pkg)`; git calls never take the index lock | |
| (`GIT_OPTIONAL_LOCKS=0`). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 13. Measured facts and pitfalls (this p150b, tt-metal 44d6650 + ETH patch) | |
| Device results of the ttaw suites and probe P2 are recorded in `common/probes/p2-trace-runner.md`. In short: | |
| - ETH dispatch opens 12x10 with 1 and 2 CQs, WORKER 11x10, `auto` -> ETH (`tests/device/test_open_device.py`). | |
| - In-trace `ttnn.copy` into persistent float32 / uint32 / bfloat16 tensors, ping-pong traces, RT-dev scalar params, | |
| packed readback, 2CQ uploads (direct and staged), CQ1 readback and replay == eager are bit-exact. | |
| - A ROW_MAJOR tensor is one page per row, and the RM `reshape` / `pad` / `concat` programs stage whole pages in L1: | |
| `reshape_rm_program_factory.cpp` needs 2 x (destination page + 80 B) per kernel copy (two copies when the pages | |
| divide each other) when the pages are not 16-byte aligned, out of ~1.43 MB free L1 per core | |
| (`TT_FATAL ... RM reshape dest staging does not fit in L1: need at least 1497632 B dest + 512 B source, have | |
| 1461248 B`, YOLOX's `[1, 1, 1, 187200]` fp32 row). Keep RM rows (last dim x element size) well below 0.5 MB; | |
| `pack_outputs` does (section 3). | |
| - A program-cache miss inside a capture raises with the op name and leaves the device usable. | |
| - numpy uint32 arrays (including values >= 2**31) upload exactly; int64 arrays must not be passed to | |
| `ttnn.from_torch` directly (they would go through bf16): use `tensors.to_host_tensor`. | |
| - `numpy.asarray(host_buffer)` fails and `numpy.from_dlpack` is read-only; `torch.from_dlpack` is a writable alias | |
| (`HostStaging`). | |
| - `ttnn.matmul` with an explicit HiFi4 + fp32 config: max |error| 0.03 vs fp64 on a 64x512x128 product; LoFi 2.4. | |
| - Runner cost (34-program toy graph, 4 MiB input): `replay` b2b equals a raw `execute_trace` (0.857 ms); streaming | |
| `run()` per frame 1.07 ms (1CQ), 1.09 ms (direct 2CQ: CQ1 must wait for the previous replay before overwriting the | |
| input), **0.92 ms with `stage_inputs=True`** (upload hidden behind the replay). Use staging when uploads are large | |
| and frames are pipelined; it buys nothing for a single synchronous request. | |
| - `ttnn.to_torch` of a host tensor copies, so `read()` results stay valid after later reads. | |
| - Review of 0.2.0 (ttaw 0.3.0, `logs/ttaw/review_*`): `ttnn.reshape` of a ROW_MAJOR tensor returns a view on the | |
| same buffer, and `ttnn.deallocate(view)` (default `force=True`) frees the base tensor's memory | |
| (`review_force_dealloc_probe.log`); a program first enqueued after a capture is a "program_cache" buffer the | |
| allocation tracker refuses before the next replay (`review_old_staged_tracking_probe.log`). Ping-pong states | |
| written through `ctx.write_target` (`output_tensor=`), D16 bank switches between two streams, `reset_state` from | |
| a bank, `write_input` followed by a CQ1 upload, and a 2CQ run after a partially failed capture are bit-exact on | |
| the p150b in 1CQ and 2CQ; with `TT_METAL_TRACE_ALLOC_TRACKING=1` bank switches and the 2CQ staged path allocate | |
| nothing after capture. | |
| - Where the 13 specs' other needs live: fp32 row gathers are `ttnn.gather` (FLOAT32, TILE; bit-exact but ~1 us per | |
| picked index), bf16 rows `ttnn.embedding` (3-5 us; probe P13, `probes/g1-dispatch-genericop-state.md`); top-k, | |
| argmax, conv / pool, attention / grid_sample facts are in `probes/g2..g4-*.md`; the op wrappers C17-C23 and | |
| kernels K1-K11 are added by the first port that needs them (PLAN.md 1.2-1.3). RT-dev values of any shape are | |
| `add_param` (re-uploaded only when they change) or, when they change every frame (index tables), `add_input`; | |
| shape buckets are variants sharing `add_output` buffers; trace segments around a host fallback are variants handed | |
| off through states or a read + `run(inputs=...)`; serve profiles pin `<ENV>_VARIANT` / `_NUM_CQS` / `_DISPATCH`, | |
| checked by `server.smoke.check_pinned`; temporal state per stream is `StreamBanks` (section 3). | |
| - Device tests: run only your own files. `common/tests/device/` also holds the probe agent's `test_probe_p*.py`, | |
| some of which spawn device subprocesses and must not run inside another pytest process. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 14. Image pre-processing (C14, `ttaw.image`) | |
| Bit-exact host emulations of the resize kernels the Autoware camera nodes deploy, driven by per-source-size lookup | |
| tables computed once and cached (a camera keeps its size). numpy only. A standard resize in their place moves the | |
| network input enough to fail the PCC gates (YOLOX: `cv2.resize` drops the det-output PCC to 0.94, | |
| research/yolox/SPEC.md section 3), so each preset is tested against a scalar transliteration of the deployed code. | |
| ```python | |
| from .ttaw.image import yolox_letterbox, yolox_letterbox_geometry | |
| bgr = rgb[:, :, ::-1] # Autoware feeds BGR8 (cv_bridge toCvCopy BGR8) | |
| x_u8, geom = yolox_letterbox(bgr) # (960, 960, 3) uint8 HWC, BGR, pad 114 | |
| x_onnx = yolox_letterbox(bgr, layout="nchw")[0].astype("float32") # == the ONNX input `images` (1, 3, 960, 960) | |
| geom.scale, geom.resized_hw, geom.mask_hw # float32 scale (decoder), (r_h, r_w), mask crop | |
| from .ttaw.image_linear import sceneseg_preprocess, sceneseg_resize, opencv_resize_nearest | |
| u8 = sceneseg_resize(bgr) # (320, 640, 3) uint8: cv::resize INTER_LINEAR, exact | |
| x = sceneseg_preprocess(bgr, channel_order="bgr") # == the SceneSeg ONNX input (1, 3, 320, 640) float32 | |
| mask_src = opencv_resize_nearest(mask_320x640, bgr.shape[:2]) # the node's INTER_NEAREST back to the source size | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `yolox_letterbox(image, dst_hw=(960, 960), *, pad_value=114, layout="hwc", out=None)` | autoware_tensorrt_yolox `preprocess.cu:43-129` @ 9ceaccf: inverted half-magnitude "bilinear" weights, `fmaf` single roundings, float -> int truncation between the passes, `lroundf`, top-left letterbox (114), channel order kept -> `(uint8 HWC or (1, C, H, W), LetterboxGeometry)`. Needs a source of at least 2x2 | | |
| | `yolox_letterbox_geometry(src_h, src_w, dst_h=960, dst_w=960)` -> `LetterboxGeometry` | `scale` (exact float32 value as a float; `scale_f32`), `resized_hw` = `(r_h, r_w)`, `mask_hw` (Autoware's mask crop, the same pair), `src_hw`, `dst_hw`, `pad_value`, `to_dict()` | | |
| | `yolox_letterbox_lut(src_h, src_w, ...)` -> `YoloxLetterboxLUT` | the cached tables (`rows`, `row_w`, `cols`, `col_w`, `k`); `.apply(image, *, out=None, chunk_rows=96)` (any channel count) | | |
| | `LUTCache(maxsize=8)` | the thread-safe LRU the presets use (`.get(key, make)`, `.hits`, `.misses`, `.clear()`) | | |
| | `lroundf(x)`, `f32(x)` | C `lroundf` (half away from zero) and one float32 rounding, for presets and their tests | | |
| | `bevdet_nearest_crop(image, *, channels="rgb", src_hw=(900, 1600), dst_hw=(256, 704), crop_hw=(140, 0), out=None)` | the `bevdet::Preprocess` plugin of autoware_tensorrt_bevdet (bevdet_vendor 0.2.1, research/bevdet/SPEC.md 3.2): nearest resize by r = (float)dst_w / src_w = 0.44f and crop (140, 0) of the 1600x900 image, `roundf(i / r + crop_h / r)` rows (318, 320, 323, ..., 898) and columns (0, 2, 5, ..., 1598) in float32 -> `(3, 256, 704)` uint8 **B, G, R** planes. `channels` is the input order (`"rgb"` from `ttaw.io.load_image`, `"bgr"` from OpenCV). The node's `cv::resize` of other image sizes to 1600x900 is the caller's (OpenCV) | | |
| | `bevdet_nearest_lut(src_hw, dst_hw, crop_hw)` -> `NearestCropLUT` | the cached gather (`rows`, `cols`, `resize`); `.apply(image, *, channels, out)`, `.apply_planes(planes [..., C, H, W])` | | |
| | `bevdet_normalize(crop, mean=BEVDET_MEAN, std=BEVDET_STD)` | `(x - mean[c]) / std[c]` in float32 on B, G, R planes (`BEVDET_MEAN` / `BEVDET_STD` are the "RGB" ImageNet numbers, applied to BGR planes as trained) | | |
| | `bevformer_preprocess(images, *, mean=BEVFORMER_MEAN_BGR, std=BEVFORMER_STD, scale=0.8, pad_divisor=32, out=None, workers=1)` | autoware_tensorrt_bevformer's pipeline (research/bevformer/SPEC.md 3.1): N BGR uint8 images of one size (the node's 1600x900) -> normalise first (`(x - mean_c) / std_c` as OpenCV's `convertTo`: BGR means 103.53 / 116.28 / 123.675, std 1) -> `cv::resize` x0.8 INTER_LINEAR **on float** -> zero pad bottom / right to /32 -> `[N, 3, 736, 1280]` float32 B, G, R planes (the ONNX input `image` without its leading 1). `workers` threads over the cameras. The node's uint8 `cv::resize` of other sizes to 1600x900 is the caller's (OpenCV) | | |
| | `bevformer_input_geometry(src_hw=(900, 1600), scale=0.8, pad_divisor=32)` | `{"src_hw", "resized_hw", "padded_hw"}`: `(int)(src * 0.8f)` in float32 and the /32 pad: (720, 1280) and (736, 1280) deployed (the network normalises its UV by the padded size) | | |
| | `bevformer_normalize(image, mean, std)` | the normalisation alone, HWC float32 | | |
| | `opencv_linear_lut(src_hw, dst_hw, *, strict=True)` -> `LinearResizeLUT`; `opencv_resize_linear_f32(image, dst_hw, *, strict=True)` | OpenCV's float INTER_LINEAR (`resize.cpp` coefficients `f = (float)((d + 0.5) * scale - 0.5)`, border clamps; each pass `fma(x1 - x0, w, x0)` in float32, horizontal first), cached tables (`rows0`, `rows1`, `wy`, `cols0`, `cols1`, `wx`); `.apply(image HWC)`, `.apply_planes(planes [..., H, W])` (strided block slices for periodic taps). `strict=True` refuses geometries not verified bit-exact (weights not multiples of 1/8: OpenCV computes other scales' coefficients differently) and exact 2x down-scales (OpenCV's INTER_AREA path) | | |
| | `PRESETS` | name -> (function, description) of the implemented presets | | |
| | `image_area.meteor_resize(image, *, out=None)` | METEOR (tier4/METEOR `hf/onnx_smoke_test.py:42-44`, research/meteor/SPEC.md 3): one uint8 frame (any channel order) -> 432x768 with OpenCV `INTER_AREA` (a copy when already 432x768). Down-scaling only (raises `ValueError` for any axis smaller than the target) | | |
| | `image_area.opencv_resize_area_u8(image, dst_hw, *, out=None)`, `opencv_area_lut(src_hw, dst_hw)` -> `AreaResizeLUT` | OpenCV `cv::resize(INTER_AREA)` of uint8 HxW / HxWxC images, any down-scale: `resizeArea_<uchar, float>` (`computeResizeAreaTab` float weights from double, horizontal then vertical float32 accumulation in table order, `cvRound`) and `resizeAreaFast_` for integer factors (float32 `s * (1.f / area)`, except the 2x2 `fast_mode` of 1-, 3- and 4-channel images: `(s + 2) >> 2`); cached tables (`fast`, `rows`, `row_w`, `cols`, `col_w`) | | |
| | `image_area.scale_intrinsics(K, src_hw, dst_hw=(432, 768))` | K of the resized image: row 0 x dst_w / src_w, row 1 x dst_h / src_h (METEOR `extract_gt.py:128-130`; anisotropic resizes keep fx != fy), float64 | | |
| | `image_linear.sceneseg_preprocess(image_bgr, *, channel_order="bgr")` | VisionPilot SceneSeg (`middleware_recipes/common/backends/onnx_runtime_backend.cpp:41-60` @ vision_pilot 04fa3e80; research/sceneseg/SPEC.md 3): one HxWx3 B, G, R uint8 frame -> the ONNX input `float32 [1, 3, 320, 640]`, bit-exact = `sceneseg_normalize(sceneseg_resize(image))`. `channel_order`: `"bgr"` (deployed: B, G, R planes, BGR-ordered ImageNet constants) or `"rgb"` (training order: R, G, B planes, RGB constants; a channel permutation of the bgr input) | | |
| | `image_linear.sceneseg_resize(image_bgr, *, dst_hw=(320, 640), out=None)`, `sceneseg_normalize(resized_bgr, *, channel_order="bgr")` | the node's squash `cv::resize(img, Size(640, 320))` (INTER_LINEAR, aspect not kept) -> uint8 `(320, 640, 3)`, channel order kept (the TT device input); the normalisation as OpenCV computes it: `convertTo(CV_32F, 1/255)` (float32 product), `cv::subtract(Scalar)` in float32, `cv::divide(Scalar)` in double rounded to float32 (a float32 division differs by 1 ulp in ~27 % of the values), `cv::split` -> NCHW. `SCENESEG_INPUT_HW`, `SCENESEG_MEAN_RGB`, `SCENESEG_STD_RGB`, `SCENESEG_CHANNEL_ORDERS` | | |
| | `image_linear.opencv_resize_linear_u8(image, dst_hw, *, out=None)`, `opencv_linear_u8_lut(src_hw, dst_hw)` -> `LinearResizeU8LUT` | OpenCV `cv::resize(INTER_LINEAR)` of uint8 HxW / HxWxC images, any scale: `resize()`'s float coefficients (`fx = (float)((dx + 0.5) * scale - 0.5)`, x-border fix only), `cvRound(w * 2048)` weights, the exact int32 horizontal pass and the uchar fixed-point vertical pass `(((b0 * (h0 >> 4)) >> 16) + ((b1 * (h1 >> 4)) >> 16) + 2) >> 2`; a copy for equal sizes and OpenCV's INTER_AREA fast path (`image_area`) for an exact 2x down-scale. Cached tables (`mode`, `cols0`, `cols1`, `a0`, `a1`, `rows`, `r0`, `r1`, `b0`, `b1`); `.apply(image, *, out=None)` | | |
| | `image_linear.opencv_resize_nearest(image, dst_hw, *, out=None)`, `opencv_nearest_lut(src_hw, dst_hw)` -> `NearestResizeLUT` | OpenCV `cv::resize(INTER_NEAREST)` (`resizeNN`: `min(floor(x * (1.0 / inv_scale)), W - 1)` in double) of HxW / HxWxC images of any dtype (a gather): e.g. SceneSeg's 320x640 mask back to the source size (`run_model_node.cpp:117-188`) | | |
| | `image_triangle.streampetr_preprocess(image, *, preset="autoware_main", channels="rgb", dst_hw=(480, 640), fma=False, out=None)` | autoware_camera_streampetr (universe main @ 9ceaccf, `lib/network/preprocess_kernel.cu:73-176` + `camera_data_store.cpp:287-316`; research/streampetr/SPEC.md 3.1): one uint8 `(H, W, 3)` frame (`channels` = its byte order) -> the float32 network input `(3, 480, 640)`: `resize = max(480 / H, 640 / W)` (float32), the virtual resized image `((int)(H * resize), (int)(W * resize))`, top rows cropped (bottom kept) and the width centred, a PIL-style triangle filter with adaptive support (`centre = (r + 0.5) * scale - 0.5`, taps `ceilf(centre - support) .. floorf(centre + support)`, `w = max(0, 1 - |d| / support)` per axis, `w = wx * wy`), accumulated **y outer, x inner** in float32, divided by the weight sum, then `(v - mean[c]) / std[c]`, planar. `preset` (`STREAMPETR_PRESETS`): `autoware_main` (RGB planes, RGB stats; PLAN D7 default), `autoware_0.52` (an rgb8 camera on 0.52.1: RGB planes, BGR-ordered stats), `awml_training` (BGR planes, BGR stats). `fma=True` models nvcc's default `--fmad=true` contraction (`fmaf` centre, channel and weight sums); `False` (default) rounds every C operation once, as the SPEC's port and the research goldens | | |
| | `image_triangle.streampetr_preprocess_batch(images, *, preset, channels, dst_hw, fma, workers=1, out=None)` | the cameras of one frame -> `(N, 3, 480, 640)` float32 (each camera its own size); `workers` threads over the cameras | | |
| | `image_triangle.autoware_resize_geometry(src_h, src_w, dst_h=480, dst_w=640)` -> `TriangleGeometry`, `triangle_resize_lut(src_hw, dst_hw=(480, 640), *, geometry=None, fma=False)` -> `TriangleResizeLUT` | `calculate_image_processing_params` in float32 (`src_hw`, `resized_hw`, `roi_hw`, `roi_start` (y, x), `resize`, `to_dict()`); the cached per-size tap tables (`rows`, `row_w`, `cols`, `col_w`; zero-weight padding) with `.apply(image u8 HxWxC) -> (C, roi_h, roi_w)` float32 weighted means (before the normalisation; any channel count). `geometry=` serves other nodes of the kernel family (BEVFusion's camera branch: verify its kernel first) | | |
| | `image_triangle.fma_f32(a, b, c)` | `fmaf` on float32 arrays: one rounding, exact (TwoSum-corrected float64; a float64 sum landing on a float32 tie is broken towards the exact value) | | |
| Facts: 0 mismatches against the scalar kernel on 12 source sizes and on the research golden input of the Autoware | |
| test image; 40-170 ms per frame on the shared host (1080p about 130 ms), a few ms per new size for the tables. The | |
| BEVDet gather equals the research golden network input of the nuScenes key-frame `sample0` bit for bit | |
| (`tests/host/test_image_bevdet_host.py`). The BEVFormer pipeline equals OpenCV 4.8.1 and 4.11.0 bit for bit on | |
| arbitrary float images for every x0.8 down-scale (the textbook two-product interpolation differs in ~10 % of the | |
| samples), the research reference's OpenCV pipeline on random uint8 images, and the research golden `image` of the | |
| nuScenes key-frame `sample0` (`tests/host/test_image_bevformer_host.py`, which also holds a scalar exact-rational | |
| oracle); 0.5 s per 6 x 1600x900 frame single-threaded on the shared host, 0.15 s with `workers=4` (OpenCV itself: | |
| 0.16 s). The device version (K9, uint8 upload + on-device resample) is an optimization item. Presets still to add | |
| (their ports): BEVFusion triangle resize with truncation (PLAN.md C14; `ttaw.image_triangle` holds the same kernel | |
| family). The METEOR `INTER_AREA` preset (`ttaw.image_area`) equals `cv2.resize` of OpenCV 4.8.1 and 4.11.0 bit | |
| for bit at every METEOR source size tested and a scalar transliteration of `resize.cpp` on small geometries | |
| (`tests/host/test_image_area_host.py`); about 0.1-0.15 s per 1920x1080 frame, 0.4 s at 2880x1860, numpy single thread. | |
| The SceneSeg preset (`ttaw.image_linear`, 0.16.0) equals `cv2.resize` INTER_LINEAR / INTER_NEAREST and OpenCV's | |
| `subtract` / `divide` scalar arithmetic of OpenCV 4.8.1 and 4.11.0 bit for bit on 23 fixed and 40 random geometries | |
| (down / up, odd sizes, 1-4 channels, the exact-2x and copy shortcuts), a scalar transliteration of `resize.cpp`, the | |
| 23 research device inputs of the SceneSeg public-data set (`input_u8_bgr_640x320.png`) and, within one ulp, the SPEC | |
| reference's float pre-processing (`tests/host/test_image_sceneseg_host.py`); 25-35 ms per 640x320 squash (414x727 to | |
| 1080p) and about 8 ms for the normalisation on the shared host, numpy single thread (OpenCV: 2-3 ms). | |
| The StreamPETR preset (`ttaw.image_triangle`, 0.17.0) equals a scalar transliteration of the CUDA kernel bit for bit | |
| in both contraction models (4 small geometries: down / up-scaling, crops, identity; 300 sampled pixels + the corners of | |
| a 1920x1080 -> 480x640 frame), a scalar transliteration of `calculate_image_processing_params` on 15 source sizes, the | |
| SPEC's numpy port (`research/streampetr/scripts/sp_common.py`) and the sha256 of the float32 network inputs of the | |
| research goldens of the PandaSet ship sample (`tests/host/test_image_triangle_host.py`). The two contraction models | |
| differ in ~43 % of the inputs by at most ~3e-5 (normalised units; the centre's last ulp moves the tap weights), far | |
| below the device's bf16 input rounding. About 0.25-0.4 s per 1920x1080 camera on the shared host (numpy, one | |
| thread), ~1.1 s per five-camera frame with `workers=4`. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 15. LiDAR host pipeline (C10 `ttaw.geometry`, C11 `ttaw.pointcloud`, C12 `ttaw.voxelize`, C15 `ttaw.nms`, C16 `ttaw.decode`) | |
| Host pre- and post-processing of the Autoware LiDAR detectors, built by the CenterPoint port (research/centerpoint/ | |
| SPEC.md sections 3 and 5) for CP, PP, TF and BF. numpy only. Each reproduces the deployed C++ / CUDA, including its | |
| float32 arithmetic where that decides a cell or a threshold; the non-deterministic parts of Autoware (time-seeded | |
| point shuffle, atomic slot / pillar order, unstable sorts) become fixed, documented policies (PLAN.md D1). | |
| ```python | |
| from .ttaw.pointcloud import StreamDensifiers, densify_sweeps, nonzero_rows | |
| from .ttaw.voxelize import PillarGridSpec, assign_pillars, decorate_pillar_features, canvas_gather_index | |
| from .ttaw.decode import CenterHeadDecodeConfig, decode_centerhead, sort_by_score, to_detected_objects | |
| from .ttaw.nms import circle_nms, iou_bev_nms, ClassRemapper | |
| spec = PillarGridSpec((-76.8, -76.8, -4.0), (76.8, 76.8, 6.0), (0.32, 0.32, 10.0), 32, 40000, order="flipped_x") | |
| cache = StreamDensifiers(num_past_frames=1).get("front") # Autoware's PointCloudDensification, per stream | |
| cache.enqueue(xyz, stamp_s, T_world_from_ego) # False: no pose -> frame skipped, not cached | |
| pts, info = cache.sweep_points() # (N, 4) x, y, z, time_lag; current sweep first | |
| pil = assign_pillars(pts, spec, shuffle_seed=0) # first 32 per cell, ids in flipped-x order | |
| feats = decorate_pillar_features(pil, spec, encoder_in_feature_size=9) # (40000, 32, 9) generateFeatures_kernel | |
| index = canvas_gather_index(pil.coords, pil.num_pillars, spec, sentinel=40000) # NHWC canvas = table[index] | |
| boxes = sort_by_score(decode_centerhead(heads, cfg)) # cfg = CenterHeadDecodeConfig.create(...) | |
| keep = circle_nms(boxes.x, boxes.y, 0.5) | |
| objs = to_detected_objects(boxes.take(keep), cfg.class_names) # ObjectClassification labels, yaw_ros | |
| objs = objs.take(iou_bev_nms(objs.x, objs.y, objs.length, objs.width, objs.yaw, objs.label)) | |
| objs.label = ClassRemapper.from_params(remapper_yaml_params).apply(objs.label, objs.length, objs.width) | |
| ``` | |
| | module / name | meaning | | |
| |---|---| | |
| | `geometry.compose`, `invert_rigid`, `transform_points`, `as_transform`, `transform_from_xyz_rpy`, `rotation_from_rpy` | float64 rigid transforms (`T_a_from_b`); `as_transform` accepts every `io.parse_transform` spelling | | |
| | `geometry.invert_affine_f32`, `compose_f32`, `transform_points_f32` | Autoware's float32 path: `Eigen::Affine3f` inverse (cofactors, column-0 determinant) and product, and `generateSweepPoints_kernel` (`((m0 x + m1 y) + m2 z) + m3`, no FMA) | | |
| | `geometry.mmdet_yaw_to_ros`, `ros_yaw_to_mmdet`, `quaternion_wxyz_from_yaw`, `yaw_from_quaternion_wxyz`, `yaw_from_rotation`, `wrap_angle` | yaw conventions (`yaw_ros = -yaw_net - pi/2`, float result as `ros_utils.cpp:55`) | | |
| | `pointcloud.SweepDensifier(num_past_frames=1, cloud_capacity=2_000_000)` | one stream's cache: `.enqueue(xyz, stamp_s, T_world_from_sensor=None, *, intensity=None) -> bool` (pose required when `num_past_frames > 0`; a missing pose leaves the cache unchanged), `.sweep_points(*, point_feature_size=4\|5) -> (points, DensifyInfo)` (float32 `world2current @ past2world` per sweep, `time_lag = float32(t_now - t_sweep)`, capacity drops a sweep and every older one), `.reset()`, `.stamps`, `.describe()` | | |
| | `pointcloud.StreamDensifiers(..., max_streams=16)` | per-`stream.id` caches with LRU eviction: `.get(id, reset=False)`, `.forget(id)`, `.streams`, `.describe()` | | |
| | `pointcloud.densify_sweeps(current_xyz, [(xyz, time_lag_s, T_current_from_sweep)], ...)` | the stateless client-side variant -> `(points, DensifyInfo)` | | |
| | `pointcloud.finite_rows`, `nonzero_rows`, `range_mask`, `seeded_permutation(n, seed)` | hygiene masks (`nonzero_rows` = `|x|+|y|+|z| > 0`, the research rule for organized-cloud fillers; `range_mask` half-open, NaN fails) and `default_rng(seed).permutation(n)` | | |
| | `voxelize.PillarGridSpec(range_min, range_max, voxel_size, max_points_per_pillar=32, max_pillars=40000, order="flipped_x"\|"raster")` | grid sizes in float32 as `centerpoint_config.hpp` (`grid_x`, `grid_y`, `num_cells`) | | |
| | `voxelize.assign_pillars(points, spec, *, shuffle_seed=None) -> PillarSet` | deterministic pillars: optional seeded permutation, half-open range, float32 `floor((v - min) / voxel)` clamped, first K per cell in input order, ids in ascending cell order capped at `max_pillars`. `PillarSet`: `points (P, K, F)`, `num_points`, `coords (z, y, x) = (0, iy, ix)`, `num_pillars`, `num_pillars_total`, `overflow`, `stats()` | | |
| | `voxelize.decorate_pillar_features(pillars, spec, *, encoder_in_feature_size=9\|10\|11)` | `generateFeatures_kernel`: raw columns, offsets to the slot-ordered float32 mean, offsets to the pillar centre `voxel/2 + coord*voxel + min`; padded slots 0 | | |
| | `voxelize.scatter_canvas(features, coords, n, spec)`, `canvas_gather_index(coords, n, spec, *, sentinel)` | `scatterFeatures_kernel` (C, H, W) and its gather form for the device (C19): cell `iy * grid_x + ix` -> pillar row or the zero sentinel row | | |
| | `decode.CenterHeadDecodeConfig.create(class_names, voxel_size_xy, range_min_xy, downsample_factor, distance_bin_upper_limits, score_thresholds, yaw_norm_thresholds, has_variance=False, has_twist=False)` | Autoware's coercions (thresholds outside [0, 1) -> 0, ascending bins, one yaw threshold per class); thresholds `(bins, classes)`; `.with_score_threshold(v)` | | |
| | `decode.decode_centerhead(heads, cfg) -> DecodedBoxes`, `sort_by_score` | `generateBoxes3D_kernel` in float32 (first-max label, no +0.5 offset, bins, `score < thr` drops, yaw-norm gate, dims (w, l, h), `atan2`), rows in cell order; stable descending sort (ties by cell). No max-pool, no top-K | | |
| | `decode.head_values_at(heads, cells) -> {name: (C, K)}`, `decode_cells(values, cells, grid_w, cfg) -> DecodedCells` | the same decoder at given cells, dropping none (agreement metrics: device vs fp32 reference at the same cells). `DecodedCells`: the `DecodedBoxes` fields (bit-identical to `decode_centerhead`'s row for every cell it keeps) plus `yaw_norm`, `distance_bin`, `score_threshold`, `yaw_norm_threshold`, `passes_score`, `passes_yaw_norm`, `.keep`, `.take`, `.to_boxes()`; `head_values_at` stores the maps at a cell set (compact goldens) | | |
| | `decode.to_detected_objects(boxes, class_names, *, has_twist=False) -> DetectedObjects` | `box3DToDetectedObject`: ObjectClassification labels (`AUTOWARE_LABELS`, `LABEL_IDS`, `semantic_label`), `yaw_ros`, (length, width, height), `orientation_availability` SIGN_UNKNOWN for car-like labels, object-frame twist; `.take`, `.boxes_xyzlwh_yaw()` (the `outputs.Detections3D` layout), `.to_records()` | | |
| | `nms.circle_nms(x, y, dist)` | `circleNMS` on score-sorted boxes: float32 `dx^2 + dy^2 < dist^2` (strict), greedy, class-agnostic | | |
| | `nms.iou_bev_nms(x, y, length, width, yaw, labels, *, search_distance_2d=10.0, iou_threshold=0.1, scores=None, sort=False)` | `perception_utils::IouBevNms::apply`: every earlier object (kept or not) suppresses, pedestrian-vs-other pairs skipped, `<=` search gate, break at `iou > thr`, keep iff `max_iou <= thr`; exact rotated-rectangle IoU (`iou_bev`, `clip_convex`, `polygon_area`, `bev_box_corners`) with the 1e-6 / 0.01 area guards | | |
| | `nms.ClassRemapper.from_params(yaml_params)` / `.from_lists(allow, min, max)` | `DetectionClassRemapper`: `.apply(labels, length, width)` with `bev_area = float(l * w)`, first allowed destination in label order | | |
| Facts (CenterPoint port, research/centerpoint): voxel tensors bit-identical to the research reference (input order | |
| and seed 0, including a 41,918-pillar overflow frame); decode + NMS + remapper give the same objects as two | |
| independent research implementations on 16 head-map goldens; about 0.3 s voxelization and 30-120 ms | |
| post-processing per 200k-point frame on the shared host (numpy). Host tests: | |
| `tests/host/test_geometry_pointcloud_host.py`, `test_voxelize_host.py`, `test_decode_nms_host.py`. Still to add (their | |
| ports): PointPainting 12-feature decoration and raster order use, TransFusion / BEVFusion decoders, OpenPCDet and | |
| METEOR NMS variants (PLAN.md C12, C15, C16). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 16. CNN ops: conv builders (C17 `ttaw.ops.conv`), up-sampling and interpolation (C18 `ttaw.ops.upsample`) | |
| Built by the YOLOX port (PLAN.md C17 / C18; probes P5, P14). A feature map is a `FeatureMap(tensor, batch, height, | |
| width, channels)`: the device tensor holds N*H*W rows of C channels (ttnn's flattened `[1, 1, N*H*W, C]`, TILE or | |
| ROW_MAJOR), plus the spatial size ttnn's layout no longer carries. Every builder prepares its device weights on its | |
| first (eager) call and reuses them, so call each layer once before capturing (the `TraceRunner` warm-up does): a | |
| host weight write inside a capture fails ("Writes are not supported during trace capture"). | |
| ```python | |
| from .ttaw.ops.conv import Conv2d, KSplitConv, ConvTranspose2d, FeatureMap, concat_channels, residual_add | |
| from .ttaw.ops.upsample import upsample_nearest, Resize2d, interp_matrix | |
| conv = Conv2d(w_oihw, b, stride=2, activation="relu6") # padding "same" (k // 2); HiFi4 + fp32 + L1 acc | |
| x = FeatureMap(ctx["image_bf16"], 1, 960, 960, 3) # or fmap_from_numpy(nchw, device) (under devrun) | |
| y = conv(x) # FeatureMap [1, 1, 480*480, 32], DRAM, TILE | |
| y = residual_add(conv2(y), y, "relu6") # RELU / RELU6 fused into ttnn.add (P5) | |
| z = concat_channels([upsample_nearest(lateral, 2), d4]) # nearest x2: bit-exact; concat along C | |
| head = Conv2d(w, b, output_dtype="float32") # fp32 logits for thresholds | |
| ks = KSplitConv(w_full, b, parts=(128, 128, 128), activation="relu") # conv(concat) without the concat (P14) | |
| up = ConvTranspose2d(w_iokk, b, stride=4, activation="relu") # k = 2: conv_transpose2d; k >= 3: linear + d2s (P5) | |
| rs = Resize2d((16, 44), (64, 176), channels=256, mode="linear", align_corners=True) # FPN_LSS-style bilinear | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `Conv2d(weight, bias=None, *, stride=1, padding="same", dilation=1, groups=1, activation=None, precision=CONV_PRECISION, output_dtype="bfloat16", output_layout="tile", slicing="auto", output_memory_config="dram", conv_config=None, weight_terms=1, name="conv")` | `ttnn.conv2d` with weights prepared once per input spec. `__call__(fmap)` (or a raw tensor + `batch=`, `height=`, `width=`) -> `FeatureMap`. `slicing="auto"`: a DRAM input runs ttnn's DRAM-sliced path with automatic slice counts (P14: traceable; explicit counts are often rejected) and writes DRAM interleaved; an L1 input runs the L1 path; `"l1"` = `Conv2dL1FullSliceConfig`; `("height"\|"width", n)` explicit (diagnostics). 1x1 stride-1 convs are matmuls on the L1 path; `output_memory_config` (`"dram"` default, `"l1"`, None = the op's sharded output) places their output. `conv_config`: extra `ttnn.Conv2dConfig` fields (optimization knobs: `act_block_h_override`, `shard_layout`, double buffers, ...). `is_matmul`, `out_hw(h, w)`, `takes_l1_path(t)`, `prepared_specs`, `release()`, `describe()` | | |
| | `weight_terms=2` | the weight and the bias as two bf16 terms (`hi = bf16(x)`, `lo = bf16(x - hi)`: ~16 significant bits; `host_terms()`), run as two convs into fp32 outputs, summed in fp32 with the activation fused into the add, cast to `output_dtype` (TILE outputs only): ~fp32 weights and bias for 2x the MACs and 3 more programs. **bf16 rounding of weights and biases, not of activations, dominates a deep CNN's error** (a per-channel bias error is the same at every pixel): YOLOX's mask agreement on PandaSet side cameras is 97.5 % with bf16 weights and 99.3 % with two terms (CPU emulation; the device matches it). The device's float32 weight path (`precision=":w=float32"`) is no substitute: 1.5x less error than bf16 vs 3.5x (fp32 output) for two terms; with a bf16 output two terms land within 1.05x of the output-rounding floor (`logs/yolox/m3_weight_precision_probe.log`, `m3_c17_weight_terms_device_r2.log`) | | |
| | `CONV_PRECISION` | `"HiFi4+fp32+l1acc"`, the default of every builder. **Without packer L1 accumulation a conv whose inner dim spans several blocks keeps its partial sums in the output dtype (bf16) between blocks** (`conv2d_op_program_factory_common.cpp:175-178`): YOLOX 3x3 256 -> 256 PCC 0.99998 / max abs 0.25 vs 0.9999993 / 0.025 with it, at the same speed (`logs/ttaw/c17_c18_device_results*.json`) | | |
| | `normalize_activation`, `unary_with_param`, `apply_activation`, `FUSED_ACTIVATIONS`, `BINARY_FUSED_ACTIVATIONS` | fused conv activations: relu, relu6, silu, gelu (exact), sigmoid; binary fused: relu, relu6. **HARDSWISH is refused** (fused, it is silently skipped: P5); apply `ttnn.hardswish` as its own op | | |
| | `KSplitConv(weight, bias, parts, *, activation=None, output_dtype="bfloat16", accumulate_dtype="bfloat16", **conv_kwargs)` | `conv(concat(x_1..x_n))` as `sum_i conv_i(x_i)` (bias on the first part, activation fused into the last add or after it); `__call__([fmaps])`. `accumulate_dtype="float32"` keeps partials fp32. `split_k_weights(w, parts)` is the host split | | |
| | `ConvTranspose2d(weight [Cin, Cout, k, k], bias=None, *, stride=k, activation=None, precision=CONV_PRECISION, output_dtype="bfloat16", method="auto", output_layout="tile", memory_config="dram")` | kernel == stride, no padding. `auto`: k = 2 -> `ttnn.conv_transpose2d`, k >= 3 -> `ttnn.linear` (bias + activation fused) + depth-to-space (RM reshape / permute / reshape); `linear_weights()` gives the (i, j, c)-ordered host matrix | | |
| | `merge_sibling_convs(weights, biases) -> (w, b, sizes)`, `split_channels(fmap, sizes)` | one conv for sibling convs reading the same input (exact), split by `ttnn.slice` on C | | |
| | `concat_channels(fmaps, *, memory_config="dram")`, `residual_add(a, b, activation=None)` | glue on maps of one spatial size | | |
| | `FeatureMap` (`rows`, `hw`, `with_tensor(t, channels=None)`, `deallocate()`), `fmap_from_numpy(nchw, device, dtype, layout, memory_config)`, `fmap_to_numpy(fmap)`, `conv_padding`, `conv_out_hw`, `is_dram` | helpers (TILE uploads touch the device: run under `bin/devrun`) | | |
| | `upsample_nearest(fmap, scale=2, *, output_layout="tile", memory_config="dram")` | integer-factor nearest (floor): TILE -> RM -> `[N, H, W, C]` -> `ttnn.upsample` -> `[1, 1, N*sH*H*sW*W, C]` -> TILE. Bit-exact (ONNX Resize nearest / asymmetric / floor) | | |
| | `interp_matrix(n_in, n_out, *, mode="linear"\|"nearest", align_corners=False, scale=None)` -> `[n_out, n_in]` float64; `resize_matrices(in_hw, out_hw, ...)` | 1-D interpolation matrices equal to `F.interpolate` (torch's half-pixel rule clamped at 0, `align_corners`, the `scale_factor` rule); host tests compare them with torch at 1e-12 | | |
| | `Resize2d(in_hw, out_hw, *, channels, batch=1, mode="linear", align_corners=False, scale=None, dtype="float32", output_dtype="bfloat16", precision="accurate", output_layout="tile", tail="auto")` | separable resize on the device for any sizes: W pass `[N*H, C, W] @ A_w^T` (transposes around one matmul), H pass `kron(I_N, A_h) @ [N*H, W'*C]`; `tail` (0.23.1): `tile` reshapes the H-pass result in TILE layout, `rm` via ROW_MAJOR (small DRAM pages: deadlocks Blackhole under ETH dispatch, SYS-1419), `auto` = `tile` for destination pages < 4 KB unless the tt-metal tree has the reshape patch (`ttaw.device.reshape_patch_present`, 0.23.2: then `rm`, the patched single-kernel path) (env `TTAW_RESIZE2D_TAIL`); fp32 operands by default (bf16 cannot hold weights like 511/1023; a device fp32 matmul is TF32-like, P12). `reference(nchw)` is its host float64 twin. `ttnn.upsample`'s own bilinear is half-pixel only (integer scales, sharded bf16 input) | | |
| Device results (`common/tests/device/test_conv_upsample_device.py`, every case eager x2, captured with cache misses | |
| forbidden, replayed on a new input, and equal to an eager run after the capture bit for bit; log | |
| `logs/yolox/m1_c17_c18_device_r1.log` / `_r2_l1acc.log`, results `logs/ttaw/c17_c18_device_results*.json`): all 54 | |
| distinct conv shapes of YOLOX seg16 (1x1 to 9x9, stride 1 / 2, 960^2 down to 30^2, fused RELU6, fp32 head preds) | |
| PCC >= 0.9999 (>= 0.9999992 for kxk with `CONV_PRECISION`); K-split 7 x 64 -> 128 0.99999; ConvTranspose k = 2 / 4 | |
| 0.99999; the 7 YOLOX nearest x2 shapes bit-exact (0.03-3.6 ms: the RM round trip dominates at 16 x 480^2); `Resize2d` | |
| align_corners True / False, integer and non-integer, PCC >= 0.9999998. Host tests: `tests/host/test_conv_upsample_host.py` | |
| (54) on the fake ttnn plus `tests/host/fake_ttnn_cnn.py` (conv2d / conv_transpose2d / upsample / matmul / argmax / | |
| where / ... for CNN graph tests; `install(fake)` returns its undo). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 17. Attention (C20, `ttaw.ops.attention`) | |
| Built by the Diffusion Planner port (PLAN.md C20; probe P7). `ttnn.transformer.scaled_dot_product_attention` has | |
| four traps; `sdpa` is the only way the ports call it: | |
| 1. `is_causal` defaults to True: with Sq == Sk it silently runs causal attention, with Sq != Sk it is rejected. | |
| `sdpa` always passes `is_causal=False`. | |
| 2. The default `scale` is `1/sqrt(padded head dim)` (wrong after a 16 -> 32 head padding). `scale=` is required. | |
| 3. The default 32/32 chunks are 2.5-7x slower and inaccurate over >= 32k keys. `sdpa` passes a program config from | |
| `chunk_config(sq, sk)` (the P7 table plus Sk-range rules) and refuses `program_config=None` once Sk >= 1,000. | |
| 4. **With a user mask the op masks the padded keys only at tile granularity**: key tiles past `ceil(Sk/32)` get | |
| `-inf`, but the last partial tile is read from the mask, whose tile padding is 0 after a host tilize | |
| (`sdpa/device/kernels/dataflow/reader_interleaved.cpp`, "Mask read"). Up to 31 padded keys then join the | |
| softmax (score `q . k_pad`, value `v_pad`): rel-L2 0.18 instead of 0.025 at 321 x 321 with 89 valid keys, and | |
| 0.06 for P7's masked 564 x 564 case. `sdpa` refuses a mask with `Sk % 32 != 0`: run masked attention on | |
| `aligned_keys(n)` keys (the extra keys masked with `-inf` in the mask; free, the tiles are padded anyway). | |
| Without a mask an unaligned Sk is exact (the op generates an element-wise padding mask). | |
| ```python | |
| from .ttaw.ops import attention as A | |
| q, k, v = A.split_qkv(qkv, 8) # [1, 1, S, 3*H*D] -> 3 x [1, H, S, D] (one program) | |
| q, k, v = A.split_q_kv(q_proj, kv_proj, 8) # Q from LN(x), K|V from x, same S | |
| qh = A.split_heads(q_cross, 8) # Q-only (cross-attention; K / V hoisted elsewhere) | |
| mask = A.expand_key_bias(ctx["key_row"], 576) # [1,1,1,Sk] bias row (persistent input) -> [1,1,Sq,Sk] DRAM | |
| o = A.sdpa(q, k, v, scale=1 / 32 ** 0.5, attn_mask=mask, concat_heads=True) # [1, 1, Sq, H*D] | |
| runner.run("plan", inputs={"key_row": A.key_bias_row(np.r_[valid, np.zeros(12, bool)])}) # 564 real + 12 masked | |
| wq, bq = A.pad_head_columns(w_q, b_q, 8, 16) # head dim 16 -> 32 in the weights (exact, free) | |
| wo = A.pad_head_rows(w_out, 8, 16) | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `sdpa(q, k, v, *, scale, attn_mask=None, program_config="auto", compute_kernel_config=None, fp32_acc=False, concat_heads=False, memory_config=None)` | `[B, H, Sq, D]` x `[B, Hkv, Sk, D]` (bf16 / bfp8 / bfp4 TILE interleaved, D % 32 == 0) -> `[B, H, Sq, D]`, or `[B, 1, Sq, H*D]` with `concat_heads=True` (the op's `output_concat_heads`, no extra program). Mask: additive bf16 TILE in DRAM, `[1\|B, 1\|H, Sq, Sk]`, Sk tile-aligned; `-inf` is safe | | |
| | `chunk_config(sq, sk) -> (q_chunk, k_chunk)`, `CHUNK_TABLE` | exact measured shapes (evidence string per entry), else Sk >= 16,384 q64 k1024; >= 4,096 q64 k512; >= 1,000 q64 k128; shorter q64 k64 (32 when the sequence fits one tile) | | |
| | `sdpa_program_config(device, sq, sk, *, q_chunk=None, k_chunk=None, exp_approx_mode=None)`, `sdpa_compute_config(*, fp32_acc=False, fidelity="HiFi2", approx=True)` | `SDPAProgramConfig` over the grid read from the device; the explicit compute config equal to the op default (P7: HiFi4 / exact exp do not help; `fp32_acc` halves the max error at ~1.4x the time) | | |
| | `split_heads(x, num_heads)`, `split_qkv(qkv, num_heads, *, num_kv_heads=None)`, `split_q_kv(q, kv, num_heads, *, num_kv_heads=None)`, `merge_heads(x)` | `[B, 1, S, H*D]` <-> `[B, H, S, D]` through `ttnn.experimental.nlp_create_qkv_heads` / `nlp_concat_heads` (one program each, bit-exact moves, the logical S kept). `split_q_kv` needs the same S for Q and K / V; cross-attention splits Q and K / V separately | | |
| | `key_bias_row(valid) -> [1, 1, 1, Sk]`, `key_bias(valid, sq) -> [1, 1, Sq, Sk]`, `expand_key_bias(row, sq)`, `aligned_keys(n)` | host masks (0 / `-inf`, float32; upload bf16 TILE) and the traced expansion of a persistent bias row (`ttnn.repeat`, DRAM; one program per plan, not per call) | | |
| | `padded_head_dim(d)`, `pad_head_dim(x)`, `pad_head_columns(w, b, H, D)`, `pad_head_rows(w, H, D)` | head dim padding to a multiple of 32 in the projection weights (exact: the padded output columns are 0, P7) | | |
| | `attention_reference(q, k, v, *, scale, bias=None)`, `attention_matmul(q, k, v, *, scale, attn_mask=None, compute_kernel_config=None)` | float64 numpy oracle; small-MHA fallback with two matmuls + softmax for what SDPA rejects (e.g. fp32 Q / K / V; tile-aligned Sk) | | |
| Device results (`common/tests/device/test_attention_device.py`, 27 cases through `TraceRunner`: eager x2 | |
| bit-identical, strict capture, replay on a new input set equal to an eager run bit for bit; set A nearly uniform | |
| attention, set B peaked; gates frozen in `test_attention_device.gates.json`; `logs/diffusion-planner/c20_device_r2.log`, | |
| results `logs/ttaw/attention_results.json`): PCC vs float64 0.99966-0.99983 at the DP shapes (576 x 576 and 352 x 352 | |
| masked, 352 / 321 x 564), TF self 500 x 500 head dim 16 -> 32, SP self 900 x 1,668 and BF cross 500 x 32,400 | |
| (q64 k1024, 0.46 ms); rel-L2 0.024-0.025 everywhere; split / merge / mask expansion bit-exact; the fp32 matmul | |
| fallback PCC 0.999998. Trace ms per call: DP fusion 576 x 576 masked 0.061-0.092, DiT self 352 x 352 masked 0.063, | |
| DiT cross 352 x 564 0.033. A canary (`test_raw_op_mask_padding_leak`) keeps the trap-4 evidence (raw op, 321 keys: | |
| rel-L2 0.183); when tt-metal masks partial tiles it fails and the refusal can be relaxed. Watcher bring-up of the | |
| smallest shape (`TT_METAL_WATCHER=2`) clean: `logs/diffusion-planner/c20_watcher_smoke.log`. Host tests: | |
| `tests/host/test_attention_host.py` (27) with `tests/host/fake_ttnn_attention.py` (an SDPA fake that keeps the op's | |
| defaults, so a caller forgetting `is_causal=False` / `scale` fails; head ops, repeat, matmul, softmax). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 18. LiDAR device modules: gather-form scatter (C19 `ttaw.ops.gather`), SECOND + SECONDFPN (C24 `ttaw.models.second`), CenterHead (C25 `ttaw.models.centerhead`) | |
| Built by the CenterPoint port (PLAN.md C19, C24, C25; probes P4, P5, P14) on the C17 builders of section 16, for | |
| CP, PP, TF and BF. Every module is built from BN-folded host (numpy) weights, prepares its device weights on its first | |
| (eager) call, keeps every output in DRAM (interleaved TILE; L1 residency is optimization work) and traces. | |
| ```python | |
| from .ttaw.ops.gather import check_index, scatter_rows, gather_rows, scatter_rows_numpy | |
| from .ttaw.models.second import ConvSpec, Second, SecondFPN, RowLinear, free_maps | |
| from .ttaw.models.centerhead import CenterHead, merge_heads | |
| from .ttaw.ops.conv import FeatureMap | |
| index = check_index(prep.canvas_index, num_rows=prep.num_pillars, sentinel=40000) # host, [1, H*W] uint32 | |
| backbone = Second([[ConvSpec(w, b, stride=s) for w, b, s in block] for block in blocks]) # 3x3 + ReLU | |
| neck = SecondFPN([ConvSpec(w0, b0), ConvSpec(w1, b1, stride=2, kind="conv_transpose"), | |
| ConvSpec(w2, b2, stride=4, kind="conv_transpose")]) # 1x1, k2 ConvT, k4 ConvT | |
| head = CenterHead(shared_spec, [128, 128, 128], [(name, hidden_spec, final_spec), ...]) # K-split + merged heads | |
| def forward(ctx): # a TraceRunner variant | |
| canvas, table = scatter_rows(pillar_rows, ctx["canvas_index"]) # [1,1,P,32] rows -> [1,1,H*W,32] TILE | |
| blocks = backbone(FeatureMap(canvas, 1, 480, 480, 32)) | |
| return head(neck(blocks)).tensor # one fp32 [1, 1, H*W, 32] tensor (15 channels used) | |
| maps = head.unpack(runner("default", inputs=...), 480, 480) # host: {name: (C, H, W)} | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `gather.check_index(index, *, num_rows, sentinel=None)` | host validation of a gather index -> contiguous `uint32 [1, N]`; every entry in `[0, num_rows)` or equal to `sentinel`, else `ValueError` (an out-of-range index reads outside the table on the device) | | |
| | `gather.gather_rows(index, table, *, sentinel=None, layout="tile", memory_config=None)` | `ttnn.embedding` of a UINT32 `[1, N]` index into a bf16 ROW_MAJOR `[1, 1, V, D]` table -> `[1, 1, N, D]`; with `sentinel` it is P4's fast `EmbeddingsType.PADDED` form (`padding_idx=sentinel`), which **returns the table's own row there**: the row must hold zeros | | |
| | `gather.with_zero_rows(rows, count=1)` | `[1, 1, P, D]` rows (TILE or ROW_MAJOR) -> ROW_MAJOR table `[1, 1, P + count, D]` with zero rows appended (one `ttnn.pad`) | | |
| | `gather.scatter_rows(rows, index, *, sentinel=None, layout="tile", memory_config=None) -> (canvas, table)` | the gather-form scatter: `canvas[n] = rows[index[n]]`, zeros where `index[n] == sentinel` (default `P`); static shapes, no canvas clear; P4: 0.21 ms at CP 480 x 480 | | |
| | `gather.gather_rows_numpy`, `gather.scatter_rows_numpy(rows, index, *, sentinel)` | host oracles | | |
| | `second.ConvSpec(weight, bias, stride=1, padding=None, relu=True, kind="conv"\|"conv_transpose", name="")` | one BN-folded layer: conv weight `[Cout, Cin, k, k]` (padding None = `k // 2`), ConvTranspose weight `[Cin, Cout, s, s]` (kernel == stride) | | |
| | `second.Second(blocks, *, policy=None, prefix="backbone")` | SECOND: `blocks[i]` = list of 3x3 `ConvSpec` (block stride on the first); `__call__(fmap) -> [block outputs]` (input kept, inner layer outputs freed), `out_shapes(h, w)`, `release()`, `describe()`. C17 `Conv2d` layers: a DRAM input runs ttnn's automatic DRAM slicing (P14) | | |
| | `second.SecondFPN(deblocks, *, policy=None, prefix="neck")` | one deblock per block output: `"conv"` 1x1 -> `RowLinear`; ConvTranspose k = 1 -> `RowLinear`, k = 2 -> `ttnn.conv_transpose2d`, k >= 3 -> linear + depth-to-space (C17 `ConvTranspose2d`, P5); every output moved to DRAM; the concat is never built | | |
| | `second.RowLinear(weight, bias=None, *, activation=None\|"relu", precision=DEFAULT_PRECISION, output_dtype="bfloat16")` | a 1x1 conv / per-row dense layer: `ttnn.linear` on `[..., R, K]` TILE rows with an explicit compute config and `core_grid` from the device (keeps the ReLU fused), output DRAM; `weight` `[K, N]` or `[N, K, 1, 1]`; device weights uploaded on the first call | | |
| | `second.DEFAULT_PRECISION` | C17's `CONV_PRECISION` (`HiFi4+fp32+l1acc`); a `ttaw.precision.PrecisionPolicy` overrides per module (`backbone.block<i>.conv<j>`, `neck.deblock<i>`, `head.shared`, `head.hidden`, `head.out`); a precision's `activations` dtype (`:a=fp32`) is the dtype of that module's OUTPUT (bf16 by default; the K-split partial sums follow the shared conv's) | | |
| | `second.free_maps(*maps)` | deallocate feature maps / tensors that still hold their buffer (`force=False`: views are left alone, P2) | | |
| | `second.second_numpy(blocks, x_nchw)`, `second.second_fpn_numpy(deblocks, blocks_nchw)` | fp32 torch oracles (NCHW) | | |
| | `centerhead.CenterHead(shared, split, heads, *, policy=None, prefix="head", output_dtype="float32", k_split=True, accumulate_dtype="bfloat16")` | `relu(conv3x3(concat(inputs)))` as a C17 `KSplitConv` over `split` (P14: 3.6 vs 14.8 ms at CP 480 x 480, more accurate), then the merged heads: `relu(x @ [Wh_1 \| ... \| Wh_n] + bh)` and the block-diagonal `@ Wo + bo` (two `RowLinear`s) -> ONE `[1, 1, H*W, C_pad]` tensor (`C_pad` = sum of head channels rounded up to 32). `shared_forward(inputs)`, `heads_forward(shared)`, `__call__(inputs)`, `unpack(rows, h, w) -> {name: (C, H, W)}`, `slices`, `release()`, `describe()` | | |
| | `centerhead.merge_heads(heads, *, pad_to=32)` | host: `(Wh, bh, Wo, bo, slices)` of `[(name, hidden 1x1 + ReLU, final 1x1)]` (exact rewrite) | | |
| | `centerhead.centerhead_numpy(shared, heads, inputs_nchw)` | fp32 torch oracle of the literal graph (concat, conv, n heads) | | |
| Device results (`common/tests/device/test_lidar_models_device.py`, TraceRunner: eager x2, strict capture, replay on a | |
| new input, replay == eager bit for bit): the CP scatter at its real shape (40,000 rows into 480 x 480) bit-exact, | |
| sentinel rows zero; CP-shaped SECOND (4 / 6 / 6 convs, 64 / 128 / 256 channels) + SECONDFPN + CenterHead with random | |
| weights at 64 x 64 and 128 x 128 (first stride 1 and 2) PCC >= 0.999 for every block, deblock, the shared conv and | |
| every head vs the torch fp32 oracle (`logs/ttaw/c19_c24_c25_device_results.json`). The CenterPoint bundle runs the same | |
| modules with the real weights (`bundles/centerpoint-p150/PORT_LOG.md`). Host tests: | |
| `tests/host/test_lidar_models_host.py` (13) on the fake ttnn plus `tests/host/fake_ttnn_lidar.py` (`embedding` with | |
| P4's PADDED semantics, `max`, `EmbeddingsType`; `install(fake)` returns its undo). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 19. grid_sample helpers (C23, `ttaw.ops.deform`) | |
| Built by the BEVDet port (PLAN.md C23; probe P9) for AlignBEV (8 history slots warped by per-frame affine grids) | |
| and FPN_LSS's align_corners bilinear up-sampling; BEVFormer (DCN / rotate) and METEOR extend it. Grids are built on | |
| the host (numpy, float32) and uploaded as ROW_MAJOR fp32 tensors (a per-frame RT-dev value or a static input); | |
| `grid_sample` is the device call with the P9 rules checked first. | |
| ```python | |
| from .ttaw.ops import deform | |
| ring = runner.add_state("ring", shape=(8, 128, 128, 96), dtype="bfloat16", layout=ttnn.ROW_MAJOR_LAYOUT) # C % 32 | |
| grid = deform.affine_grid(transforms_8x6, (128, 128), valid=[h >= 0 for h in history]) # [8, 128, 128, 2] fp32 | |
| runner.add_param("grid", grid, shape=grid.shape, dtype="float32", layout=ttnn.ROW_MAJOR_LAYOUT) | |
| aligned = deform.grid_sample(ctx["ring"], ctx["grid"]) # in a variant: bilinear, zeros, align_corners=False | |
| up = deform.grid_sample(f2_rm, static_grid) # static_grid = deform.resize_grid((16, 16), (64, 64)) | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `pixel_to_grid(px, size)` | pixel coordinate (0 = first pixel centre) -> normalised `(2 px + 1) / size - 1` (float32): the align_corners=False form, exact for integer `px` when `size` is a power of two | | |
| | `kernel_pixels(g, size)` | what the device kernel recovers: `g * size/2 + (size - 1)/2` in float32 (`grid_sample_reader_common.hpp`) | | |
| | `affine_pixel_coords(transforms, out_hw)` | `ix = a w + b h + c`, `iy = d w + e h + f` per batch (float32, one rounding per op): the BEVDet AlignBEV pixel affine | | |
| | `affine_grid(transforms [N, 6], in_hw, out_hw=None, *, valid=None)` | fp32 grid `[N, H, W, 2]`; `valid[n] == False` -> `OUTSIDE` (an exact zero sample: an empty history slot) | | |
| | `resize_pixel_coords(in, out, align_corners=True)`, `resize_grid(in_hw, out_hw, *, align_corners=True, batch=1)` | static bilinear-resize grids (float64 source pixels, align_corners=False normalisation); for align_corners=True every source pixel is inside, so it equals `F.interpolate` up to the weight truncation | | |
| | `padded_channels(c)`, `pad_channels(x)` | the input channel rule (multiple of 32; pad channels stay exactly 0 through the op) | | |
| | `emulate_grid_sample(x, grid)` | host emulation of the kernel (fp32 fractions, weights **truncated** to bf16, taps outside skipped): the device test's oracle | | |
| | `grid_sample(x, grid, *, compute_kernel_config=None, memory_config=None, batch_output_channels=False)` | `ttnn.grid_sample` bilinear / zeros / align_corners=False; refuses a TILE input or grid, input C % 32 != 0, a grid batch != input batch, a bf16 grid, and a FLOAT32 input with `fp32_dest_acc_en` (P9 defect); output DRAM by default | | |
| | `OUTSIDE` | `-4.0`: a normalised coordinate whose four taps miss every input | | |
| Rules (P9): fp32 grids for every data-dependent grid (bf16 grids: PCC 0.9988 at 53 px, 0.992 at 128 px); C padded to | |
| a multiple of 32; grids expanded per batch (no broadcast); ROW_MAJOR interleaved / height-sharded input and grid; | |
| never a FLOAT32 input with `fp32_dest_acc_en`; call the op with align_corners=False on the align_corners=False | |
| normalisation of the pixel coordinates you mean (zero padding acts in pixel space, so it is the same sampling as the | |
| model's align_corners=True formulation, but integer coordinates stay exact: an identity warp is a bit-exact copy; | |
| the literal ac1 normalisation leaves 2.8-7.6 % of values off by up to 1/32). The bilinear weights are truncated to | |
| bf16 (0.33-0.6 % rrmse floor; K5 removes it); C18 `Resize2d` (fp32 interpolation matmuls) is the accurate | |
| alternative for static resizes. | |
| Device results (`tests/device/test_deform_device.py`, TraceRunner: eager warm-up, strict capture, replay on new | |
| data, replay == eager bit for bit; `logs/ttaw/deform_device_results.json`): AlignBEV [8, 128, 128, 96] identity slots | |
| bit-exact, `OUTSIDE` slots exactly 0, pad channels 0, rigid-motion slots PCC >= 0.9999 vs torch fp32 and within 2 | |
| bf16 ulps of `emulate_grid_sample`; the FPN_LSS resizes 16 -> 64 (640 ch) and 64 -> 128 (512 ch) PCC >= 0.99999 vs | |
| `F.interpolate(align_corners=True)`, corners exact. Host tests: `tests/host/test_deform_host.py` (13). The BEVDet | |
| bundle runs the same calls at its real shapes (`bundles/bevdet-p150/PORT_LOG.md`: AlignBEV moving slots PCC 0.999995 | |
| vs the CPU reference). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 20. ResNet builders (C27, `ttaw.models.resnet`) | |
| Built by the BEVDet port (PLAN.md C27) for its ResNet-50 image backbone (6 cameras as batch) and CustomResNet BEV | |
| encoder; BEVFormer and METEOR reuse it. Built on the C17 builders of section 16 from BN-folded host weights; every op | |
| traces, each conv prepares its device weights on its first (eager) call, every activation is a DRAM-interleaved TILE | |
| `[1, 1, N*H*W, C]` map and intermediates are freed as soon as they are consumed. | |
| ```python | |
| from .ttaw.models import resnet as rn | |
| prov = lambda module: rn.ConvParams(w[module], b[module], stride, padding_tblr, dilation, groups) # BN folded | |
| stem = rn.ResNetStem(prov, "img_backbone.conv1") # 7x7 s2 + ReLU + maxpool 3x3 s2 | |
| r50 = rn.ResNetStages(prov, rn.mmdet_names("img_backbone"), [3, 4, 6, 3]) # bottleneck, pytorch style | |
| c4, c5 = r50(stem(img_fmap), keep=(2, 3), free_input=True) # FPN inputs; the rest freed | |
| bev = rn.ResNetStages(prov, rn.custom_resnet_names("img_bev_encoder_backbone"), [2, 2, 2], block="basic", | |
| factory=rn.conv_factory("HiFi4+fp32+l1acc:w=fp32")) | |
| f0, f1, f2 = bev(bev_in, keep=(0, 1, 2)) | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `ConvParams(weight [Cout, Cin/g, kh, kw], bias=None, stride=(1, 1), padding=(t, b, l, r), dilation=(1, 1), groups=1)` | one BN-folded conv as a provider returns it (numpy float32); a provider is `module name -> ConvParams` (BEVDet's reads `OnnxWeights` by consuming node, so no kernel size or stride is re-typed) | | |
| | `conv_factory(precision=DEFAULT_PRECISION, **conv_kwargs)` | `(module, params, activation) -> callable`: one C17 `Conv2d` per module; `precision` a spec / `Precision` or a `PrecisionPolicy` resolved per module name. **DCN hook:** pass your own factory and return a deformable layer for the modules you name | | |
| | `mmdet_names(prefix)`, `custom_resnet_names(prefix)` | `(stage, block) -> {conv1, conv2[, conv3], downsample, block}`: `layer{s+1}.{b}.conv{k}` / `.downsample.0`, or `layers.{s}.{b}.conv{k}` / `.downsample` | | |
| | `Bottleneck(provider, names, *, has_downsample, factory)` | `relu(conv3(relu(conv2(relu(conv1(x))))) + idt)`, stride on `conv2` (pytorch style) | | |
| | `BasicBlock(provider, names, *, has_downsample, factory)` | `relu(conv2(relu(conv1(x))) + idt)`; the projection may be any conv (CustomResNet: 3x3 stride 2) | | |
| | `ResNetStem(provider, conv_module, *, factory=None, pool=True)` | `maxpool3x3s2p1(relu(conv1(x)))`; the pool runs DRAM-sliced and writes TILE | | |
| | `ResNetStages(provider, names, layers, *, block="bottleneck"\|"basic", factory=None, downsample_first=())` | `__call__(x, keep=(), *, free_input=False) -> [kept stage outputs]` (input kept unless `free_input`; every other block output freed once consumed); `stages`, `convs()` | | |
| | `maxpool_out_hw(h, w)`, `DEFAULT_PRECISION` (= C17 `CONV_PRECISION`) | helpers | | |
| | `resnet_numpy(x_nchw, provider, names, layers, *, block, stem=None, keep=(), downsample_first=())` | fp32 torch oracle of the same blocks | | |
| Precision: the ports pick it. BEVDet runs `HiFi4+fp32+l1acc:w=fp32` (fp32 weights and biases: the bf16 rounding of | |
| the BN-folded biases, applied per channel at every pixel, cost it depth PCC 0.9998 vs 0.99993 and 3 of the 52 sample0 | |
| boxes, at no measured time cost); C17's `weight_terms=2` is the other accurate option. Device results | |
| (`tests/device/test_resnet_device.py`, TraceRunner protocol; `logs/ttaw/resnet_device_results.json`): stem + two | |
| bottleneck stages on [6, 3, 128, 352] and a CustomResNet stage (3x3 s2 projection) on [1, 864, 128, 128] -> 160, | |
| random weights, PCC >= 0.999 vs `resnet_numpy` with both precisions, replay == eager bit for bit. Host tests: | |
| `tests/host/test_resnet_host.py` (6, fake ttnn + `fake_ttnn_cnn` + a local max-pool fake). The BEVDet bundle runs the | |
| real ResNet-50 / CustomResNet (`bundles/bevdet-p150/PORT_LOG.md`: depth / feat PCC 0.99994, chained heads >= 0.9995 | |
| vs the fp32 CPU reference). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 21. Query heads: top-k (C21 `ttaw.ops.topk`), heatmap peaks (C22 `ttaw.ops.heatmap`), TransFusion head (C26 `ttaw.models.transfusion_head`) | |
| Built by the TransFusion port (PLAN.md C21, C22, C26; probes P6, P7, P8, P13) for TF, BF and PT: everything after the | |
| dense heatmap -- the 3x3 local maximum, the top-K proposal selection and the query decoder -- as one traceable device | |
| graph with one packed output. numpy only at import. | |
| ```python | |
| from .ttaw.ops.heatmap import LocalMax, sigmoid_heat | |
| from .ttaw.ops.topk import TopKSelect, selection_metrics | |
| from .ttaw.models.transfusion_head import QueryHeadWeights, TransFusionQueryHead, assemble_transfusion | |
| local_max = LocalMax(192, 192, 5, rule="transfusion") # "bevfusion" / "ptv3": pooled_classes, passing cells | |
| head = TransFusionQueryHead(QueryHeadWeights(...), prefix="decoder") # host numpy weights, BN folded | |
| head.prepare(device) # tables + LayerNorm rows (before any capture) | |
| def forward(ctx): # a TraceRunner variant | |
| heat, _ = sigmoid_heat(logits_fp32) # fp32 sigmoid -> bf16 (max_pool2d / top-k are bf16) | |
| heat_nms = local_max(heat) # [1, 1, H*W, C] TILE | |
| out = head(lidar_feat, heat_nms, ctx.get("topk_indices")) # None: device top-k; else teacher forcing | |
| return pack_outputs(out) # heads (K x 32 fp32), query heat (Kp x 32), indices | |
| raw = runner("default", inputs=...) | |
| outs = assemble_transfusion(head.unpack_heads(raw["heads"]), raw["query_heat"], raw["indices"], bev_pos, | |
| num_classes=5, num_proposals=500, cells=36864) # cls_score0 / bbox_pred0 / dir_cls_pred0 | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `heatmap.local_max_masks(h, w, classes, *, rule="transfusion"\|"bevfusion"\|"ptv3", pooled_classes=None, kernel=3)` | the constant masks `(m_pass, m_eq)` (NHWC rows `[h*w, classes]`, 0 / 1, disjoint) of `keep = m_pass + (heat == maxpool_kxk_s1_p(k//2)(heat)) * m_eq`: TF interior `m_eq`, no pass; BF pooled classes (default 0-3) as TF, the others pass everywhere; PT `m_pass = 1 - m_eq` (border and unpooled classes pass) | | |
| | `heatmap.LocalMax(h, w, classes, *, rule, pooled_classes=None, kernel=3, memory_config="dram")` | the device chain on a bf16 TILE `[1, 1, h*w, classes]` heat (`max_pool2d` p1, `eq`, `multiply` by `m_eq` or `addcmul(m_pass, eq, m_eq)`, `multiply` by the heat); masks uploaded by the first call; `.numpy(heat_rows)` its host twin; refuses fp32 (max_pool2d is bf16 only) | | |
| | `heatmap.sigmoid_heat(logits, *, dtype="bfloat16") -> (heat, sig)` | the sigmoid in the logits' dtype (fp32 logits: fp32 sigmoid), cast once; `sigmoid_heat_numpy` its host model | | |
| | `heatmap.local_max_numpy`, `local_max_literal_numpy(heat_nchw, rule)` | oracles: the masked chain, and each model's literal `p0` + border / unpooled-class form | | |
| | `topk.TopKSelect(cells, classes, k)` | `__call__(heat_nhwc) -> (indices, pos, cls)` (uint32 ROW_MAJOR `[1, padded_k(k)]`): `class_major_row` (transpose + untilize + reshape to one bf16 row, index `c * cells + p` = the ONNX `TopK` order), `topk_indices` (`topk_large_indices`, descending, ties unspecified), `decode` | | |
| | `topk.decode_class_major(indices, cells, classes, *, clamp=True)` | exact decode with fp32 comparisons (`cls = sum_c idx >= c * cells`, `pos = idx - cls * cells`; no division), clamped so a 0xFFFFFFFF sentinel cannot become an out-of-range gather index; uint32 ROW_MAJOR rows for `ttnn.embedding` | | |
| | `topk.padded_k(k)`, `K_MULTIPLE`, `MASK_VALUE` (-1e30), `SENTINEL` | k rounded up to a multiple of 16 (500 -> 512); mask scores with a finite value, never `-inf` (P8) | | |
| | `topk.topk_numpy(scores, k)`, `class_major_numpy`, `decode_class_major_numpy`, `selection_metrics(dev_idx, ref_idx, *, reference_scores=None, flat_scores=None, strong=0.1)` | ONNX-semantics host top-k (ties to the lower index), layouts, and the set-based metrics of a device selection (`shared`, `strong_total` / `strong_kept` / `strong_overlap` / `strong_missed`) | | |
| | `topk.local_max_margins(indices, heat, height, width, *, kernel=3)` (0.21.0) | per class-major proposal, the relative margin `(heat[p] - max same-class neighbour) / heat[p]` in the reference heat BEFORE the local max (1.0 without neighbours): the decisive / near-tie split of a strong-proposal gate (TransFusion gates margins > 0.5 %, reports the rest with `near_tie_report`) | | |
| | `topk.topk_with_values(scores_tile, k)` | the bf16 `ttnn.topk` composite (values + indices, P8) | | |
| | `transfusion_head.QueryHeadWeights(height, width, num_classes, num_proposals, class_table, bev_pos, self_posembed, cross_posembed, self_attn, cross_attn, norms, ffn, heads)` | model-agnostic host record (`MHAWeights`, `PosEmbedMLP`, `PredictionHead`; `[in, out]` layouts): `.validate()`, `.query_pos_table()` / `.key_pos_table()` (the position MLPs on `bev_pos`, float64), `.merged_heads()` | | |
| | `transfusion_head.TransFusionQueryHead(weights, *, policy=None, prefix="head", attn_fp32_acc=False)` | `__call__(lidar_feat, heat_nms, indices=None, taps=None) -> {"heads", "query_heat", "indices"}`; pieces `init_queries`, `keys`, `decoder`, `predict`, `query_heat`; `unpack_heads(raw) -> {name: (c, K)}`; `prepare(device)`, `release()`, `describe()` | | |
| | `transfusion_head.QueryHeadOracle(weights)` | the literal head in torch fp32 (posembed MLPs on `bev_pos[pos]`, unpadded attention, six heads) from NCHW maps and given indices: the device tests' reference | | |
| | `transfusion_head.assemble_transfusion(heads, query_heat, indices, bev_pos, *, num_classes, num_proposals, cells)` | TF's (mmdeploy-patched) outputs on the host in float32: `cls_score0 = query_heat * sigmoid(heatmap)`, `bbox_pred0 = [center + query_pos, height, dim, vel]`, `dir_cls_pred0 = rot` | | |
| Device form of the head (every rewrite exact in real arithmetic; the oracle runs the literal forms): queries | |
| `lidar_feat[pos] + class_table[cls]` (bf16 `ttnn.embedding` row gathers, uint32 indices from the exact decode); the | |
| query position embedding is a function of `pos` only, so it is a constant **QPE table** `self_posembed(bev_pos)` | |
| gathered by `pos` (`query_pos` itself, `k + 0.5` up to 191.5, never passes through bf16: S:transfusion:241); keys | |
| `lidar_feat + KPE` with the constant `KPE = cross_posembed(bev_pos)`; head dim 16 -> 32 in the projection weights | |
| (C20 `pad_head_columns` / `pad_head_rows`, exact), `scale` from the weights, C20 `sdpa` (non-causal, chunk table); | |
| self-attention on the logical K queries with no mask (Q / K / V from one fused projection); LayerNorm with the | |
| model's epsilon (`ttnn.layer_norm(o, residual_input_tensor=x, epsilon=...)`); merged heads (`relu(x @ [Wh_i]) @ | |
| blockdiag(Wo_i)`, fp32 out). One decoder layer only (a second layer would re-embed the predicted centres: no constant | |
| table). Precision `HiFi4+fp32:w=fp32` (`default_policy()`; per module `<prefix>.self_attn` / `.cross_attn` / `.ffn` / | |
| `.heads` / `.norm`). | |
| Device results (`tests/device/test_transfusion_head_device.py`, 13 cases through `TraceRunner`: eager x2, strict | |
| capture, replay on a new input set, replay == eager bit for bit; watcher bring-up of the small shapes clean; | |
| `logs/transfusion/m1/j2_ttaw_{small_watcher,full}.log`, results `logs/ttaw/c21_c22_c26_device_results.json`, gates | |
| frozen in `test_transfusion_head_device.gates.json`): the masked local max is **bit-exact** against each model's | |
| literal form at TF 192x192x5, BF 180x180x5 and PT 256x256x7 (plateaus and bright borders: ties everywhere); the | |
| top-512 of TF 184,320 / BF 162,000 / PT 458,752 values has exactly the host's value multiset, unique in-range | |
| indices in descending order, and an exact decode (also exhaustively over all 184,320 TF indices); surviving `-inf` | |
| scores give 0xFFFFFFFF (22 of 32 in the canary) and the decode clamps them in range; the query head with random | |
| weights, teacher-forced proposals, vs the literal fp32 oracle: TF (36,864 keys) PCC >= 0.9970 (queries 0.999998, | |
| decoder layers >= 0.9996, heads >= 0.9970), BF (32,400 keys, the (i, j) `bev_pos`) >= 0.9922, query heat exact. | |
| The TransFusion bundle runs the same modules with the deployed weights (`bundles/transfusion-p150/PORT_LOG.md`). | |
| The selection chain at TF size (sigmoid, local max, class-major row, top-512, decode, one row gather) replays in | |
| about 1.1 ms (`logs/transfusion/m1/exp1_full_*.log`: 2.0-2.4 ms for two chains). Host tests: | |
| `tests/host/test_transfusion_head_host.py` (29) with `tests/host/fake_ttnn_query.py` (`topk_large_indices` with | |
| HIGHEST-index ties, bf16-only `max_pool2d`, `layer_norm` with ttnn's 1e-12 default epsilon, `eq` / `ge` / | |
| `subtract` / `addcmul`, an `embedding` that also records computed indices inside a capture). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 22. Pillar feature net and input staging (C24 companion, `ttaw.models.pillars`) | |
| The PillarFeatureNet of the pillar detectors (CenterPoint, PointPainting, TransFusion) on the device and the host | |
| staging of its two inputs. Promoted by the PointPainting port (ttaw 0.19.0) from the CenterPoint bundle's | |
| `tt/pfn.py` + `FrameStaging` (device-verified there: PFN PCC 0.99997) with the TransFusion bundle's generalisation | |
| (any hidden / output width; pillars of K < 32 points). ttaw 0.21.0 adds TransFusion's two-term PFN (`terms=2`) and | |
| the TransFusion bundle runs on this module from then on; the CenterPoint bundle keeps its own copy until it re-vendors | |
| and switches (its PORT_LOG notes it). | |
| ```python | |
| from .ttaw.models.pillars import FrameStaging, PfnPlan, TtPillarFeatureNet, features_shape, pack_inputs | |
| from .ttaw.ops.gather import scatter_rows | |
| plan = PfnPlan.from_weights(w0, b0, w1, b1) # BN-folded [in, out] matrices: (F, H), (2H, O) | |
| pfn = TtPillarFeatureNet(plan, num_pillars=40000, precision="HiFi4+fp32:w=fp32:a=fp32") # a= : hidden h / z dtype | |
| runner.add_input("features", shape=features_shape(40000), dtype="bfloat16") # ROW_MAJOR, 1 KB / pillar | |
| runner.add_input("canvas_index", init=np.full((1, H * W), 40000, np.uint32), dtype="uint32") | |
| def forward(ctx): # a TraceRunner variant | |
| rows = pfn(ctx["features"]) # [1, 1, P, O] bf16 ROW_MAJOR | |
| canvas, table = scatter_rows(rows, ctx["canvas_index"], sentinel=40000) # C19: NHWC [1, 1, H*W, O] TILE | |
| ... | |
| staging = FrameStaging(40000, H * W) # persistent host tensors, written in place per frame | |
| out = runner("default", inputs=staging.write(features, canvas_index, num_pillars)) | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `PfnPlan.from_weights(w0, b0, w1, b1, *, in_padded=16)` | the device matrices of the identity-folded split PFN: `W0 (16, H_pad)`, `Wz = [I_H \| W1a \| 0] (H_pad, Z)`, `Wy = [W1b ; I_O ; 0] (Z, O)` (`H_pad` = H rounded up to 32, `Z` = H + O rounded up to 32); F <= 16, O % 32 == 0 (the C19 gather table). `.forward_numpy(features, *, hidden_round=None)`: its float32 host emulation (`hidden_round` models the hidden dtype) | | |
| | `pfn_reference_numpy(features, w0, b0, w1, b1)` | the literal encoder in float64 (concat form; max over the given slots, padded slots included): the oracle | | |
| | `TtPillarFeatureNet(plan, *, num_pillars, precision="HiFi4+fp32:w=fp32", name="pfn")` | `[1, 1, P, 32*16]` bf16 ROW_MAJOR -> `[1, 1, P, O]` bf16 ROW_MAJOR rows: reshape to `[1, 1, P*32, 16]`, tilize, `linear` W0 + ReLU, `linear` Wz, `max` over the 32 slots (dim -2), `linear` Wy + b1 + ReLU, untilize. HiFi4 required (the identity blocks); `Precision.activations` = the dtype of the hidden `h` / `z` (the output rows are always bf16); `.release()`, `.describe()` | | |
| | `TtPillarFeatureNet(..., terms=2, bias_add=True)` (0.21.0), `TWO_TERM_PRECISION = "HiFi4+fp32:w=bf16"` | the two-term PFN (TransFusion's default): every linear a `TwoTermLinear` (weights / biases as bf16 `hi + lo`, products `hi@W_hi + hi@W_lo + lo@W_hi` with fp32 outputs), fp32 `h` / `z` / slot max split by `split_terms` before the next linear, bf16 output rows; `bias_add=True`: bias-free `RowMatmul` (`ttnn.matmul`, fp32 out) + one exact fp32 `ttnn.add` per bias (`ttnn.linear`'s fused bias rounds an fp32 output to ~11 bits on this tree); needs `w=bf16` | | |
| | `split_terms(t)`, `RowMatmul(weight, *, precision, name)`, `TwoTermLinear(weight, bias=None, *, activation=None, precision=TWO_TERM_PRECISION, name, bias_add=True)` | the two-term building blocks (`split_terms`: fp32 -> bf16 `hi`, `lo = bf16(t - hi)`; also used by TransFusion's two-term SECOND convs) | | |
| | `features_shape(P)` -> `(1, 1, P, 512)`; `pack_features(features (P_kept, K, F), P, *, out=None)` | the device input: one 1 KB row per pillar (32 slots x 16 features, zero-padded), rows past `P_kept` zero; **K < 32 slots: slots K..31 repeat slot 0** (TransFusion's 20-point pillars: the slot maxima are unchanged, zero rows would add `relu(b0)`) | | |
| | `pack_inputs(features, canvas_index, num_pillars, *, capacity)` | `{"features", "canvas_index"}` as numpy (tests and tools); the index range-checked (C19 `check_index`, sentinel = `capacity`) | | |
| | `FrameStaging(num_pillars, num_cells)` | persistent bf16 `features` / uint32 `canvas_index` host tensors (`tensors.HostStaging`); `.write(features, canvas_index, num_pillars=None)` converts only the kept pillars (torch fp32 -> bf16 RNE, K < 32 slot rule), zeroes the rows the previous frame used beyond them, copies the checked index and returns the `TraceRunner` inputs; `.zero_copy` (False: falls back to `pack_features` + `HostStaging.write`, same values) | | |
| Device form (exact rewrites, also in floating point): one pillar = one 32-row tile, so a slot max is a tile-column | |
| reduction; `relu(. + c)` is monotonic, so the per-pillar term `c = m @ W1b + b1` is added after the max (no slot | |
| broadcast); `h` and the max ride through the matmuls in identity blocks (x * 1.0 is exact with HiFi4). Every pillar | |
| row is computed, also past the frame's pillar count: the scatter index never points there, so no count reaches the | |
| device. Precision: the bf16 rounding of the hidden `h` / `z` is the PFN's largest error term (PointPainting's CPU | |
| emulation on its golden frames: output rel. L2 0.0045 with bf16 hidden tensors, 0.0020 with fp32 hidden tensors and | |
| TF32 weights, 0.0017 for the bf16 output rounding alone), hence the `a=fp32` option; a model with large-magnitude input | |
| columns (absolute x / y up to 121.6 m) can rewrite its first layer around the pillar centre on the host (PointPainting | |
| `tt/pfn_input.py`, a promotion candidate) and feed this module unchanged. | |
| Device results (`tests/device/test_pillars_device.py`, TraceRunner: eager x2, strict capture, replay on input set B, | |
| replay == eager bit for bit; the 512-pillar cases first under `TT_METAL_WATCHER=2`, clean; random weights, bf16 | |
| inputs, vs the literal float64 encoder on the same inputs; `logs/pointpainting/m1/j1_pillars_*.log`, results | |
| `logs/ttaw/pillars_device_results.json`): CenterPoint widths (9 -> 16, 32 -> 32) PCC 0.999996 / rel. L2 0.0025; | |
| PointPainting (16 -> 16, 32 -> 32) bf16 hidden 0.999996 / 0.0025, fp32 hidden 0.9999975 / 0.0024; TransFusion (11 -> | |
| 32, 64 -> 64, 20-point pillars) bf16 hidden 0.999995 / 0.0027, fp32 hidden 0.999997 / 0.0025; the real capacity | |
| (40,000 pillars, 28,137 kept, fp32 hidden) 0.9999975; `FrameStaging`'s zero-copy path equals `pack_features` + bf16 | |
| RNE for a large frame, a smaller one (its stale rows zeroed) and 20-point pillars. The PointPainting bundle runs it | |
| with its real weights (PFN PCC 0.99999 on two golden frames, `bundles/pointpainting-p150/PORT_LOG.md`). Host tests: | |
| `tests/host/test_pillars_host.py` (12; fake ttnn: the plan exact against the literal encoder for the CP / PP / TF | |
| widths and the 20-slot rule, packing, the module through `TraceRunner` replay == eager with bf16 and fp32 hidden | |
| tensors, `FrameStaging`'s fallback path). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 23. Segment reductions: K1 `segment_reduce` and the log-step fallback (`ttaw.ops.segment`) | |
| Built by the FRNet port (PLAN.md 1.3 row K1; owned by it) for FRNet's six `ScatterElements(max)` sites (points | |
| sorted by frustum cell); PTv3 pooling and BEVPool (sum, optimization phase) are of the same form once the host sorts | |
| the rows. A *segmented* tensor holds rows sorted by segment id: segment `s` is the contiguous row range | |
| `[off[s], off[s + 1])` (CSR offsets, non-decreasing); rows past `off[-1]` are padding that nothing reads. | |
| ```python | |
| from .ttaw.ops import segment as seg | |
| off = seg.segment_offsets(seg_id_sorted, num_segments) # host CSR offsets, uint32 [S + 1] | |
| row = seg.check_offsets(off, num_segments=S, num_rows=N_cap, length=S + 1) # [1, K] for the device | |
| k1 = seg.SegmentReduce(seg.SegmentReduceSpec(num_rows=N_cap, channels=128, num_segments=M_cap + 1, | |
| input_dtype="float32", input_layout="tile", | |
| output_dtype="bfloat16")) | |
| runner.add_input("offsets", init=row, dtype="uint32") # RT-dev: rewritten per frame | |
| def forward(ctx): # a TraceRunner variant | |
| table = k1(point_rows_fp32_tile, ctx["offsets"]) # [1, 1, M_cap + 1, 128] bf16 ROW_MAJOR; empty rows -> 0 | |
| return gather_rows(ctx["pix2vox"], table, sentinel=M_cap) # C19: straight into a gather table | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `SegmentReduceSpec(num_rows, channels, num_segments, input_dtype="float32", input_layout="tile", output_dtype="float32", op="max", empty_value=0.0, workers_per_core=2)` | static shape of one call site (COMPILE). `num_rows % 32 == 0`, `channels % 32 == 0`; bf16 / fp32 in, TILE or ROW_MAJOR; bf16 (round to nearest even) / fp32 out; `op` `"max"` only on the device (`sum` / `mean` raise `NotImplementedError`: the data-movement RISC-Vs have no float unit, see below); `.l1_bytes_per_worker()`, `.check_l1(budget)` | | |
| | `SegmentReduce(spec, *, name, grid=None)` | `__call__(x, offsets, *, output=None)` -> `[1, 1, num_segments, C]` ROW_MAJOR DRAM (allocated per call, or the persistent `output` of that spec): ONE `ttnn.generic_op`, eager or inside a capture. `x` `[1, 1, num_rows, C]` DRAM interleaved of the spec's dtype / layout; `offsets` uint32 ROW_MAJOR `[1, K >= num_segments + 1]` DRAM. `program(x, offsets, out)` (the `ProgramDescriptor`), `allocate_output(device)`, `output_spec()`, `describe()`. `grid` overrides the device grid (tests: one core) | | |
| | `segment_reduce(x, offsets, *, num_segments, op="max", output_dtype=None, empty_value=0.0, output=None)` | one-shot form (spec from the tensors) | | |
| | `logstep_segment_max(x, shift_tables, last_rows, *, layout="tile")` | the exact stock-op fallback (PLAN.md 1.3): `R` rounds `x = max(x, x[idx_r])` (untilize + `ttnn.embedding` + `ttnn.maximum`: 3 programs each), then a PADDED gather of `last_rows` from the result with a zero row appended at `R` (empty segments -> 0). bf16 in / out only (`ttnn.embedding` tables are bf16); exact for segments of up to `2^R` rows; `R` is COMPILE (bucket it per frame) | | |
| | `segment_offsets(seg_id, num_segments, *, num_rows=None)`, `check_offsets(offsets, *, num_segments, num_rows, length=None)` | host CSR offsets from sorted ids (ids `>= num_segments` = trailing padding); validation (starts at 0, non-decreasing, ends `<= num_rows`) into the device `[1, length]` row | | |
| | `segment_rounds(max_len)`, `segment_shift_tables(seg_id, rounds)`, `segment_last_rows(offsets, *, empty_row)` | the fallback's tables: `ceil(log2 L)`; `idx_r[i] = i - 2^r` in the same segment, else `i`; each segment's last row (`empty_row` = the row count: the appended zero row) | | |
| | `segment_reduce_numpy(x, offsets, num_segments=None, *, op="max"\|"min"\|"sum"\|"mean", empty_value=0.0)`, `logstep_segment_max_numpy(x, shift_tables, last_rows)` | oracles (max / min exact in the input dtype; sum / mean in float64) | | |
| | `worker_segments(offsets, num_segments, num_workers, *, num_rows=None)` | the kernel's work split, host twin (tests, diagnostics) | | |
| | `OPS`, `DEVICE_OPS`, `KERNEL` | `("max", "min", "sum", "mean")`, `("max",)`, `"segment_reduce_dm.cpp"` | | |
| **Kernel** (`ops/kernels/segment_reduce_dm.cpp`, one source for both data-movement RISC-Vs of every core). Exact by | |
| construction: the maximum is taken on integer keys of the float bit patterns (`key = bits ^ ((bits >> 31) & | |
| 0x7fffffff)`: signed order == float order, -0 < +0; NaN unsupported), so bf16 / fp32 results are bit-exact, and an | |
| fp32 input with a bf16 output equals rounding first (RNE is monotonic). Work split without per-core runtime args | |
| (probe P3 rule; common args = the three buffer addresses): worker `w` (= core index x 2 + RISC-V) owns the segments | |
| whose cost `off[s] + s` (rows + one output row) lies in `[w T / W, (w + 1) T / W)`, found on the device by two binary | |
| searches over the offsets (64-byte DRAM probes), so nothing in the host tables depends on the grid and ETH (12x10) | |
| and WORKER (11x10) give identical outputs. Each worker streams its contiguous rows in units of 32 (one tile-row in | |
| TILE mode; double-buffered NOC reads), reduces into an L1 accumulator and writes one output row per segment through | |
| a ring of 8 staging rows; empty segments get `empty_value` (FRNet's frustum2pixel zero row comes out of the kernel). | |
| Offsets are clamped to `[0, num_rows]` and to non-decreasing order, so a corrupt table cannot make it read outside | |
| `x`. L1 per worker: 2 units + 1 KiB offsets window + 128 B probe + C x 4 B accumulator + 8 output rows (FRNet's | |
| widest site, 256 fp32 channels: 75 KiB). Hang protocol (PLAN.md 4.4): no CB push / pop (CBs are plain scratch), no | |
| multicast, no semaphores; DRAM reads keep the source's 64-byte alignment offset in L1 (Blackhole | |
| `NOC_DRAM_READ_ALIGNMENT_BYTES`), every read is waited on (the barrier also invalidates the Blackhole L1 cache) | |
| before use, staging rows are reused after `noc_async_writes_flushed`, and the kernel ends with both barriers. | |
| Why max only: the data-movement RISC-Vs (rv32im + Zba / Zbb, no F extension) compare integers natively (Zbb `max`) | |
| but would emulate float adds in software; a `sum` / `mean` K1 needs the compute engine (the optimization-phase | |
| consumers, PTv3 / BEVPool, add it with their device tests). | |
| Device results (`tests/device/test_segment_reduce_device.py`, the FRNet port's jobs; results | |
| `logs/ttaw/k1_segment_reduce_device_results.json`): tiny bring-up under `TT_METAL_WATCHER=2` clean (64 x 32, 6 | |
| segments incl. empty / 1-row / tile-row-crossing ones; one core with 1 and 2 workers; eager x2 and a 2-call trace | |
| replayed on new data, bit-exact); FRNet's real sites through `TraceRunner` with the offsets rewritten per frame (10 | |
| replays x 5 rewrites), every output bit-exact vs the oracle and replay == eager, 240 workers (12x10 ETH grid, 2 per | |
| core): OT128 encoder site 160,000 x 256 fp32 -> 60,000 fp32 rows 5.52 ms per replay, OT128 point sites 160,000 x 128 | |
| fp32 -> 60,001 bf16 rows 2.89 ms, QT128-like dense segments (6,255 segments, one of 500 rows) 2.22 ms, 2,048-row cases | |
| 0.04-0.08 ms; the log-step fallback equals K1 bit for bit (4 rounds on 4,096 rows; 7 rounds on 120,000 rows with | |
| 111-row segments) (`logs/frnet/m1a/j1_bringup_and_graph.log`, watcher log `logs/frnet/m1a/j1_tiny_watcher.log`, | |
| `logs/ttaw/k1_segment_reduce_device_results.json`). Inside the FRNet graph (six sites per frame, OT128 sample): every | |
| tap downstream PCC >= 0.9995 vs the fp32 reference (`bundles/frnet-p150/PORT_LOG.md`). A 2,000-launch soak | |
| (`test_soak`) runs in the port's next job. Host tests: | |
| `tests/host/test_segment_host.py` with `tests/host/fake_ttnn_segment.py` (program-descriptor records, a | |
| `generic_op` that emulates the kernel from its compile-time args and checks the P3 rules and the split, `maximum`, | |
| `hardswish`, `softmax`, `addcmul`). | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 24. Sparse-conv rulebooks (C13, `ttaw.sparse`) | |
| Built by the BEVFusion port (PLAN.md C13; owned by it) for BEVFusion's 21-layer sparse encoder and PTv3's | |
| sub-manifold stages: the host side of the gather-GEMM sparse conv (C28). A sparse tensor is an active voxel set | |
| `coords (N, 3)` = `(x, y, z)` in `spatial_shape (X, Y, Z)`; every spconv layer of the Autoware graphs | |
| (`GetIndicePairsImplicitGemm` + `ImplicitGemm`, spconv 2.3.8) is a **gather**, for sub-manifold and strided layers | |
| alike: `out[o] = sum_k W[:, k, :] @ in[nmap[o, k]]` (taps with `-1` contribute nothing), `W` the KRSC filter | |
| `[Cout, k0, k1, k2, Cin]` (k0 along x), tap `k = (a * k1 + b) * k2 + c`. numpy only. | |
| ```python | |
| from .ttaw import sparse as sp | |
| l1 = sp.SparseLevel(coors_zyx[:, ::-1], (1440, 1440, 41)) # (x, y, z) rows = the feature rows | |
| nm1 = sp.subm_neighbor_map(l1, 3) # (N1, 27) int32, spconv's 13-query rule | |
| l2, d1 = sp.strided_neighbor_map(l1, 3, 2, 1) # next level + (N2, 27) gather map | |
| bucket, overflow = sp.select_capacity(l2.num_active, (40_960, 147_456, 262_144)) | |
| idx = sp.im2col_index(nm1, sentinel=cap, rows=cap) # (cap / 32, 27 * 32) uint32, tile-ordered | |
| # device (C28): ttnn.embedding(idx, table [1, 1, cap + 1, Cin_pad], layout=TILE, padding_idx=cap, PADDED) | |
| # -> experimental.view [1, 1, cap, 27 * Cin_pad] -> matmul [27 * Cin_pad, Cout] (BN folded) + ReLU | |
| bev = sp.dense_gather_index(l_out, sentinel=n_out) # (Z, X * Y): one canvas gather per z slice | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `SparseLevel(coords, spatial_shape)` | an active set: exact int64 keys over the bounding box of grid + voxels (sorted keys + `searchsorted`, no aliasing, out-of-grid voxels allowed), duplicates refused; `.lookup(q) -> rows` (-1 = none), `.in_grid(q=None)`, `.linear_index()` (`(x * Y + y) * Z + z`), `.num_active`, `.describe()` | | |
| | `subm_neighbor_map(level, kernel=3, dilation=1, *, rule="spconv"\|"exact")` | `(N, K)` int32. spconv uses `padding = (k // 2) * dilation` for sub-manifold layers. `"spconv"`: the 13-query + symmetric-write construction (`indices.py:1532,865,838-845`): a tap `k < K // 2` needs the **neighbour** inside the grid, a tap `k > K // 2` the **output voxel** inside it (S:ptv3:158); the centre tap is the voxel itself. `"exact"`: the plain lookup. Equal for in-grid voxel sets (every BEVFusion level); `"alias"` (PTv3's linear-key aliasing, S:ptv3:388) raises `NotImplementedError` until PTv3 adds it | | |
| | `strided_neighbor_map(level, kernel, stride, padding, dilation=1) -> (out_level, nmap)` | regular sparse conv: `o` active iff some active `i = o * s - p + k * d` (`0 <= o < out_shape`); outputs in ascending linear-key order; one input per (output, tap); inputs must lie inside their grid | | |
| | `conv_out_shape(shape, kernel, stride, padding, dilation=1)`, `kernel_offsets(kernel, dilation=1)`, `iter_taps(kernel)` | grid arithmetic `(in + 2p - d(k - 1) - 1) // s + 1`; tap offsets in tap order | | |
| | `gather_index(nmap, *, sentinel, rows=None)` | `(rows, K)` uint32: `-1` and capacity rows -> `sentinel` (an explicit zero row of the table: C19's PADDED rule) | | |
| | `im2col_index(nmap, *, sentinel, rows=None, block=32)` / `im2col_numpy(table, index, *, kernel_volume)` | the tile-ordered index `(rows / 32, K * 32)`, `index[b, k * 32 + v] = gather[b * 32 + v, k]`: a TILE `ttnn.embedding` of it into a `Cin % 32 == 0` table yields `[rows / 32, K * 32, Cin]` whose tiles, in memory order, are those of the `[rows, K * Cin]` im2col matrix (row `o` = its K tap rows concatenated, tap-major), so `ttnn.experimental.view` gives the matmul input (probe P4: 4.7 ms for 256k x 27 x 32); the oracle builds that matrix | | |
| | `dense_gather_index(level, *, sentinel, split_axis=2)`, `dense_numpy(features, level)` | to-dense as gathers (`index[s, a * B + b]`, one canvas per slice of `split_axis`) and the ScatterND reference `[X, Y, Z, C]` | | |
| | `select_capacity(count, buckets) -> (bucket, overflow)` | smallest bucket holding `count` (grid-independent constants, PLAN.md 0.2); the largest and `True` when none does | | |
| | `sparse_conv_numpy(features, nmap, weight, bias=None)`, `dense_conv3d_numpy(features, level, weight, out_coords, *, stride, padding, dilation)` | oracles: the gather form (float64 accumulation) and a brute-force dense 3-D cross-correlation (Conv3d semantics) at the active outputs | | |
| | `pairs_count(nmap)`, `neighbor_stats(nmap)` | rulebook size (FLOPs = 2 Cin Cout pairs), `{rows, pairs, mean_taps, max_taps}` | | |
| Facts (BEVFusion port): on the sample-rosbag frame #144 (98,334 voxels) every level's active count and every layer's | |
| pair count equal the research reference's (`research/bevfusion/scripts/bevfusion_ref.py`, an independent | |
| `index_add_` scatter form): 679,726 / 325,630 / 1,383,475 / 334,714 / 736,959 / 147,183 / 253,269 / 20,710 pairs for | |
| subm1 / down1 / subm2 / down2 / subm3 / down3 / subm4 / conv_out; all four sub-manifold and four strided maps take | |
| about 2.2 s in numpy on the shared host (a C++ / device builder is optimization work). Host tests: | |
| `tests/host/test_sparse_host.py` (15): sub-manifold and strided maps equal a dense Conv3d at BEVFusion's layer | |
| shapes, strided output sets equal the brute-force set, `"spconv"` equals a literal transcription of spconv's | |
| construction also with out-of-grid voxels (where it differs from `"exact"`), the im2col tile order equals the im2col | |
| matrix's for Cin 32 / 64 / 128, BEVFusion's level chain 1440x1440x41 -> 720x720x21 -> 360x360x11 -> 180x180x5 -> | |
| 180x180x2. C28 (the device gather-GEMM) is added with its device tests by the BEVFusion port's M1. | |
| --------------------------------------------------------------------------------------------------------------------- | |
| ## 25. Gather-GEMM sparse encoder (C28, `ttaw.models.sparse_encoder`) | |
| Built by the BEVFusion port (PLAN.md C28; owned by it; probe P4) on the C13 rulebooks of section 24: every spconv | |
| layer as ONE gather + ONE matmul on the device, any chain of sub-manifold / strided layers with basic-block residuals, | |
| traceable, no count on the device. PTv3's sub-manifold stages are the second consumer. | |
| ```python | |
| from .ttaw import sparse as sp | |
| from .ttaw.models.sparse_encoder import SparseConvSpec, SparseEncoder, dense_gather, pack_table | |
| specs = [SparseConvSpec("conv_input", w0, b0, "subm1", in_terms=2), # KRSC [Cout, k0, k1, k2, Cin], BN folded | |
| SparseConvSpec("l1.conv1", w1, b1, "subm1"), | |
| SparseConvSpec("l1.conv2", w2, b2, "subm1", residual="conv_input"), # relu(conv + b + identity) | |
| SparseConvSpec("down1", w3, b3, "down1"), ...] | |
| enc = SparseEncoder(specs, policy=policy, weight_terms=1) # fp32 weights (TF32-like, P12) | |
| runner.add_input("vfe", pack_table(vfe, cap1, terms=2), dtype="bfloat16", layout=ttnn.ROW_MAJOR_LAYOUT) | |
| runner.add_input("subm1", sp.im2col_index(nm1, sentinel=cap1, rows=cap1), dtype="uint32") # per index name | |
| def forward(ctx): | |
| y, table = enc(ctx["vfe"], {n: ctx[n] for n in enc.index_names}) # last fp32 output + its bf16 table | |
| return dense_gather(table, ctx["dense0"]) # to-dense slice [1, 1, cells, C] | |
| ``` | |
| | name | meaning | | |
| |---|---| | |
| | `SparseConvSpec(name, weight, bias, index, residual=None, relu=True, in_terms=1, out_terms=1)` | one layer: `weight` KRSC (or `[Cout, K, Cin]`), `index` = the gather index name it reads (shared by the convs of a level), `residual` = an earlier layer's name; `.cin_pad` / `.cout_pad` (`terms * C` rounded up to 32), `.matrix()` (the `[K * cin_pad, cout_pad]` im2col weight), `.bias_row()`, `.lo_mask()` | | |
| | `im2col_weight(weight3, *, cin_pad, cout_pad, in_terms, out_terms)` | host matrix: tap-major rows `k * cin_pad + j`; the `lo` channel block repeats the weights, `out_terms=2` repeats the columns | | |
| | `TtSparseConv(spec, *, precision="HiFi4+fp32:w=fp32", weight_terms=1)` | `__call__(table, index, identity=None)` -> fp32 TILE `[1, 1, rows, cout_pad]`: TILE + PADDED `ttnn.embedding` of the tile-ordered `[rows / 32, K * 32]` index into the bf16 ROW_MAJOR `[1, 1, V + 1, cin_pad]` table (row `V` zero: the sentinel), `ttnn.experimental.view` to `[1, 1, rows, K * cin_pad]`, C24 `RowMatmul` (fp32 out; `weight_terms=2`: bf16 hi + lo weights, two matmuls), one exact fp32 bias add (+ the identity, ReLU fused into the last add); `gather(table, index)`; `to_table(y)` -> the next bf16 table (zero row appended; `out_terms=2`: `[bf16(y) \| bf16(y - bf16(y))]`) | | |
| | `SparseEncoder(specs, *, policy=None, weight_terms=1, name="sparse")` | the chain (consecutive layers must agree on terms / widths, residual sources on the output layout); `__call__(table, indices, taps=None, keep_table=True) -> (y, table)`; `index_names`, `input_width`, `input_terms`; intermediate outputs freed after their last (residual) use unless tapped | | |
| | `dense_gather(table, index)` | to-dense slice: C19 PADDED gather of a `[1, cells]` C13 `dense_gather_index` row -> `[1, 1, cells, C]` TILE | | |
| | `pack_table(x, capacity, *, terms=1, width=None, out=None)`, `table_rows_numpy(x, *, terms, width)` | the first table on the host (`(1, 1, capacity + 1, width)` float32 with bf16-exact values; rows past `N` and the sentinel zero) | | |
| | `sparse_layer_numpy(table, nmap, spec, *, identity=None)`, `sparse_encoder_numpy(specs, features, maps, *, round_tables=True)` | oracles (float64 matmul); `round_tables` models the device's bf16 / two-term tables, False is the fp32 network | | |
| **Activation terms.** `ttnn.embedding` is bf16-only, so every layer's input is rounded to bf16 once. A two-term | |
| table (`in_terms=2`) carries `hi` and `lo = bf16(x - hi)` side by side in the channels (one gather, one matmul over | |
| `[hi | lo] @ [W ; W]`); where `2 C <= 32` (a VFE input of 4 columns, BEVFusion's 16-channel level 1) the `lo` block | |
| rides in channels the 32-wide table pads anyway. On BEVFusion's PandaSet frame the CPU emulation of the 21 layers | |
| gives conv_out PCC 0.999995 with bf16 tables (fp32 weights), 0.999999 with two-term tables: the encoder is not the | |
| precision bottleneck (BEVFusion PORT_LOG.md "Sparse precision"). | |
| Device results (`tests/device/test_sparse_encoder_device.py`, TraceRunner: eager x2, strict capture, replay on a | |
| second frame, replay == eager bit for bit; the small cases first under `TT_METAL_WATCHER=2`, clean; | |
| `logs/bevfusion/m1/j1_c28_small_watcher.log`, `j2_c28_bf.log`, results `logs/ttaw/c28_sparse_encoder_device_results.json`): | |
| a 6-layer BEVFusion-shaped chain (220 voxels; two-term VFE, one- / two-term tables, fp32 and two-term weights) PCC | |
| >= 0.99999 vs the oracle with the same table rounding; the 21-layer BEVFusion encoder with random weights (1440 x | |
| 1440 x 41 -> 180 x 180 x 2; 28.8 k / 37.5 k / 25.9 k / 14.0 k / 13.8 k active voxels in capacities 32,768 / 40,960 | |
| / 32,768 / 16,384 / 16,384) every layer PCC >= 0.99998, to-dense gathers bit-exact, **27.3 ms per replay** (21 | |
| gathers + matmuls + table conversions + 2 to-dense gathers; first port, not optimized). Host tests: | |
| `tests/host/test_sparse_encoder_host.py` (15; `tests/host/fake_ttnn_sparse.py`: a batch-form `embedding` and an | |
| `experimental.view` that reinterprets TILE memory tile by tile, so the tile-ordered index order is checked). | |