Download code/tt_diffusion_planner/ttaw/image.py from changh95/diffusion-planner-p150: direct link, hf CLI and curl.
- Browser
- Download file 37.4 kB
-
https://huggingface.co/changh95/diffusion-planner-p150/resolve/main/code/tt_diffusion_planner/ttaw/image.py
- Command line
-
hf download hf://changh95/diffusion-planner-p150/code/tt_diffusion_planner/ttaw/image.py
-
curl -L -o image.py https://huggingface.co/changh95/diffusion-planner-p150/resolve/main/code/tt_diffusion_planner/ttaw/image.py
37.4 kB
| # SPDX-License-Identifier: Apache-2.0 | |
| """C14 image pre-processing: bit-exact host emulations of the Autoware camera nodes' resize kernels, driven by | |
| per-source-size lookup tables (LUTs) that are computed once and cached. | |
| Each preset reproduces one deployed CUDA / OpenCV routine exactly (integer outputs equal, not "close"), because a | |
| standard resize in its place changes the network input enough to fail the PCC gates (YOLOX: ``cv2.resize`` drops the | |
| detection-output PCC to 0.94, research/yolox/SPEC.md section 3). The tables hold everything that depends only on the | |
| source and destination sizes (source indices, interpolation weights, letterbox geometry); applying them is a few | |
| vectorised numpy passes over the image. numpy only; no ttnn, no torch, no OpenCV. | |
| Presets (PLAN.md C14; the other camera ports add theirs here): | |
| ========================== ========================================================================================= | |
| ``yolox_letterbox`` autoware_tensorrt_yolox ``resize_bilinear_letterbox_nhwc_to_nchw32_batch_kernel`` | |
| (autoware_universe @ 9ceaccf, ``perception/autoware_tensorrt_yolox/src/preprocess.cu:43-129``, | |
| launched by ``tensorrt_yolox.cpp:243-332``): a "bilinear" resize with inverted, | |
| half-magnitude weights, float -> int truncation between the horizontal and vertical passes, | |
| ``lroundf``, top-left letterbox padded with 114, BGR kept, no normalisation. | |
| ``bevdet_nearest_crop`` the ``bevdet::Preprocess`` TensorRT plugin of autoware_tensorrt_bevdet (``bevdet_vendor`` | |
| 0.2.1; research/bevdet/SPEC.md section 3.2): nearest resize by r = 0.44f + crop (140, 0) | |
| of the 1600x900 image, ``roundf(i / r + crop_h / r)`` rows / columns in float32, planar | |
| B, G, R uint8 (the normalisation ``bevdet_normalize`` is separate: device-side in the TT | |
| port). The node's ``cv::resize`` of non-1600x900 images to 1600x900 stays with the caller. | |
| ``bevformer_preprocess`` autoware_tensorrt_bevformer's ``preprocessing_pipeline`` (autoware_universe @ 9ceaccf, | |
| ``src/preprocessing/normalize_multiview_image.cpp:70-95`` + | |
| ``augmentation_transforms.cpp:90-222``; research/bevformer/SPEC.md section 3.1): | |
| normalise FIRST (BGR uint8 -> float, ``(x - mean_c) / std_c``, BGR mean, std 1), then | |
| ``cv::resize`` x0.8 with INTER_LINEAR **on float data** (1600x900 -> 1280x720), then pad | |
| bottom / right with 0 (normalised space) to a multiple of 32 (1280x736), CHW. The node's | |
| ``cv::resize`` of non-1600x900 uint8 images to 1600x900 stays with the caller. | |
| ========================== ========================================================================================= | |
| Exactness notes for ``yolox_letterbox`` (all arithmetic is float32 in the kernel): | |
| - ``scale = min(dst_w / (float)src_w, dst_h / (float)src_h)``; ``r_h = (int)(scale * src_h)``, | |
| ``r_w = (int)(scale * src_w)``; the kernel receives ``k = (float)(1.0 / (double)scale)`` (``preprocess.cu:121-128``). | |
| - per output row ``h``: ``c = k * (float)(h + 0.5)``, ``src_row = clamp(lroundf(c) - 1, 0, src_h - 2)``, | |
| weight ``t = (1 + lroundf(c) - c) / 2`` (float32), and the same per column (``preprocess.cu:48-59, 77-88``). | |
| - ``lerp1d(a, b, t) = fma(t, b, fma(-t, a, a))`` with float operands (CUDA's float overload, i.e. ``fmaf``: | |
| one rounding per call); the horizontal results are truncated to int before the vertical ``lerp1d`` | |
| (``lerp1d`` takes ``int a, int b``), and the result is ``lroundf``-ed (``preprocess.cu:43-46, 57, 101-103``). | |
| - ``t`` lies in [0.25, 0.75], so ``t`` is a multiple of 2**-25 and ``a - t*a`` / ``t*b + inner`` with integer | |
| ``a, b`` in [0, 255] need at most 34 significant bits: float64 holds them exactly, and one ``astype(float32)`` is | |
| exactly the single rounding of ``fmaf``. | |
| - Rows ``h >= r_h`` and columns ``w >= r_w`` are 114 (``preprocess.cu:109-110``); the output values are integers in | |
| [0, 255], so this module returns uint8 (the network input is ``float(value)``, ``norm_factor`` 1.0). | |
| Verified against two independent scalar transliterations of the CUDA source (``common/tests/host/test_image_host.py`` | |
| and research/yolox/scripts/{check_preprocess,verify_preproc_scalar}.py: 0 mismatches). Residual uncertainty, from the | |
| SPEC: if nvcc picked the double ``fma`` overload, 0-2 of 3000 samples would differ by 1 (not testable without CUDA). | |
| Exactness notes for ``bevformer_preprocess`` (float32 data; OpenCV's float INTER_LINEAR path): | |
| - normalisation: ``Mat - mean`` / ``std`` evaluates as ``convertTo(alpha = 1/std, beta = -mean/std)`` with float | |
| ``alpha`` / ``beta`` (one fused multiply-add per value); with the deployed std 1 it is ``x - float(mean_c)``. | |
| - resize coefficients (``resize.cpp``): ``scale = 1.0 / ((double)dst / src)``, ``f = (float)((d + 0.5) * scale - | |
| 0.5)``, ``s = floor(f)``, ``w = f - s`` (float), clamped to the border (``s < 0`` -> ``s = 0, w = 0``; | |
| ``s >= src - 1`` -> ``s = src - 1, w = 0``); the destination size is ``(int)(src * 0.8f)`` in float. | |
| - each pass is ``fma(x1 - x0, w, x0)`` in float32 (the difference rounded once, then one rounding of the fused | |
| product-sum), horizontal pass first, then vertical. This is what OpenCV 4.8.1 and 4.11.0 compute on this x86-64 | |
| host (AVX2 + FMA) bit for bit for every x0.8 down-scale, on arbitrary float images (not just uint8 - mean): the | |
| plain two-product form ``x0 * (1 - w) + x1 * w`` differs from it in about 10 % of the samples by 1 ulp | |
| (``tests/host/test_image_bevformer_host.py`` holds both checks). The weights of a x0.8 resize are multiples of | |
| 1/8, so float64 holds every fused product-sum exactly and one ``astype(float32)`` is the single rounding. OpenCV | |
| computes the coefficients of other scales differently, so :func:`opencv_linear_lut` refuses them by default. | |
| """ | |
| from __future__ import annotations | |
| import threading | |
| from collections import OrderedDict | |
| from dataclasses import dataclass | |
| from typing import Any, Callable, Dict, Hashable, Optional, Tuple | |
| import numpy as np | |
| __all__ = [ | |
| "LetterboxGeometry", | |
| "YoloxLetterboxLUT", | |
| "yolox_letterbox", | |
| "yolox_letterbox_geometry", | |
| "yolox_letterbox_lut", | |
| "lroundf", | |
| "f32", | |
| "LUTCache", | |
| "PRESETS", | |
| "YOLOX_PAD_VALUE", | |
| "NearestCropLUT", | |
| "bevdet_nearest_lut", | |
| "bevdet_nearest_crop", | |
| "bevdet_normalize", | |
| "BEVDET_MEAN", | |
| "BEVDET_STD", | |
| "LinearResizeLUT", | |
| "opencv_linear_lut", | |
| "opencv_resize_linear_f32", | |
| "bevformer_input_geometry", | |
| "bevformer_normalize", | |
| "bevformer_preprocess", | |
| "BEVFORMER_MEAN_BGR", | |
| "BEVFORMER_STD", | |
| ] | |
| F32 = np.float32 | |
| YOLOX_PAD_VALUE = 114 # preprocess.cu:109-110 (letterbox fill), the YOLOX training pad value | |
| def f32(x: Any) -> np.ndarray: | |
| """``x`` rounded once to float32 (round-to-nearest-even), as a numpy float32 scalar or array.""" | |
| return np.asarray(x).astype(F32) | |
| def lroundf(x: Any) -> np.ndarray: | |
| """C ``lroundf``: round half away from zero, for float32 inputs (exact in float64). Returns int64.""" | |
| x = np.asarray(x, dtype=np.float64) | |
| return (np.sign(x) * np.floor(np.abs(x) + 0.5)).astype(np.int64) | |
| class LUTCache: | |
| """A small thread-safe LRU of tables keyed by the source / destination sizes (a camera stream keeps its size, | |
| so the table is built once per stream; Autoware re-creates its buffers only when the source size changes, | |
| ``tensorrt_yolox.cpp:253-291``).""" | |
| def __init__(self, maxsize: int = 8): | |
| self.maxsize = int(maxsize) | |
| self._items: "OrderedDict[Hashable, Any]" = OrderedDict() | |
| self._lock = threading.Lock() | |
| self.hits = 0 | |
| self.misses = 0 | |
| def get(self, key: Hashable, make: Callable[[], Any]) -> Any: | |
| with self._lock: | |
| if key in self._items: | |
| self._items.move_to_end(key) | |
| self.hits += 1 | |
| return self._items[key] | |
| value = make() # outside the lock: building a table takes a few ms | |
| with self._lock: | |
| self.misses += 1 | |
| self._items[key] = value | |
| self._items.move_to_end(key) | |
| while len(self._items) > self.maxsize: | |
| self._items.popitem(last=False) | |
| return value | |
| def clear(self) -> None: | |
| with self._lock: | |
| self._items.clear() | |
| self.hits = self.misses = 0 | |
| def __len__(self) -> int: | |
| return len(self._items) | |
| # ------------------------------------------------------------------------------------------- YOLOX letterbox | |
| class LetterboxGeometry: | |
| """Where the source image lands in the network input, and what the decoder and the mask crop use. | |
| ``scale`` is the float32 scale of ``preprocess.cu:121`` (also ``scales_`` of ``tensorrt_yolox.cpp:285``, used to | |
| map boxes back, and the mask scale of ``tensorrt_yolox.cpp:528-536``); ``resized_hw`` = ``(r_h, r_w)``, the | |
| un-letterboxed region at the top-left of the ``dst_hw`` canvas. Autoware's mask crop ``(out_h, out_w)`` = | |
| ``((int)(src_h * scale), (int)(src_w * scale))`` is the same pair (float multiplication commutes).""" | |
| src_hw: Tuple[int, int] | |
| dst_hw: Tuple[int, int] | |
| scale: float # exact float32 value, stored as a Python float | |
| resized_hw: Tuple[int, int] | |
| pad_value: int = YOLOX_PAD_VALUE | |
| def scale_f32(self) -> np.float32: | |
| return F32(self.scale) | |
| def mask_hw(self) -> Tuple[int, int]: | |
| """``(out_h, out_w)`` of Autoware's mask (``getMaskImageGpu``): the un-letterboxed region.""" | |
| return self.resized_hw | |
| def to_dict(self) -> Dict[str, Any]: | |
| return {"src_hw": list(self.src_hw), "dst_hw": list(self.dst_hw), "scale": self.scale, | |
| "scale_f32_hex": F32(self.scale).tobytes()[::-1].hex(), "resized_hw": list(self.resized_hw), | |
| "mask_hw": list(self.mask_hw), "pad_value": self.pad_value, | |
| "letterbox": "top-left (padding at the bottom / right)"} | |
| def yolox_letterbox_geometry(src_h: int, src_w: int, dst_h: int = 960, dst_w: int = 960, | |
| pad_value: int = YOLOX_PAD_VALUE) -> LetterboxGeometry: | |
| """Letterbox geometry of ``preprocess.cu:121-123`` (float32 scale, int truncation).""" | |
| src_h, src_w, dst_h, dst_w = int(src_h), int(src_w), int(dst_h), int(dst_w) | |
| if src_h < 2 or src_w < 2: | |
| raise ValueError(f"the Autoware resize needs a source of at least 2x2 pixels, got {src_w}x{src_h}") | |
| if dst_h < 1 or dst_w < 1: | |
| raise ValueError(f"bad destination size {dst_w}x{dst_h}") | |
| scale = min(F32(dst_w) / F32(src_w), F32(dst_h) / F32(src_h)) # float32 divisions | |
| scale = F32(scale) | |
| r_h = int(F32(scale * F32(src_h))) | |
| r_w = int(F32(scale * F32(src_w))) | |
| return LetterboxGeometry((src_h, src_w), (dst_h, dst_w), float(scale), (min(r_h, dst_h), min(r_w, dst_w)), | |
| int(pad_value)) | |
| def _axis_table(n_dst: int, n_src: int, k: np.float32) -> Tuple[np.ndarray, np.ndarray]: | |
| """Source index and weight of every destination row (or column): ``preprocess.cu:48-51, 77-88``.""" | |
| c = (k * (np.arange(n_dst) + 0.5).astype(F32)).astype(F32) # centroid = scale * (float)(h + 0.5) | |
| rc = lroundf(c) | |
| idx = rc - 1 | |
| idx = np.where(idx < 0, 0, idx) | |
| idx = np.where(idx >= n_src - 1, n_src - 2, idx) | |
| weight = ((F32(1) + rc.astype(F32)).astype(F32) - c).astype(F32) # 1 + lroundf(c) - c (float32) | |
| weight = (weight / F32(2)).astype(F32) | |
| return idx.astype(np.int64), weight | |
| class YoloxLetterboxLUT: | |
| """The per-size tables of the YOLOX letterbox: source row / column of every destination row / column and the | |
| float32 weights. Build with :func:`yolox_letterbox_lut` (cached); apply with :meth:`apply`.""" | |
| geometry: LetterboxGeometry | |
| rows: np.ndarray # (r_h,) int64: source row hi (hi + 1 is the second tap) | |
| row_w: np.ndarray # (r_h,) float32: vertical weight | |
| cols: np.ndarray # (r_w,) int64: source column wi | |
| col_w: np.ndarray # (r_w,) float32: horizontal weight | |
| k: float # the kernel's float32 ``scale`` argument (1 / scale) | |
| def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None, chunk_rows: int = 96) -> np.ndarray: | |
| """Letterbox one BGR uint8 image ``(src_h, src_w, 3)`` (any channel count) -> uint8 ``(dst_h, dst_w, C)`` | |
| HWC, channel order unchanged. ``out`` (uint8, that shape) is written in place when given.""" | |
| g = self.geometry | |
| img = np.asarray(image) | |
| if img.ndim != 3 or img.dtype != np.uint8 or tuple(img.shape[:2]) != g.src_hw: | |
| raise ValueError(f"expected a uint8 ({g.src_hw[0]}, {g.src_hw[1]}, C) image, got {img.dtype} {img.shape}") | |
| dst_h, dst_w = g.dst_hw | |
| r_h, r_w = g.resized_hw | |
| shape = (dst_h, dst_w, img.shape[2]) | |
| if out is None: | |
| out = np.empty(shape, np.uint8) | |
| elif out.shape != shape or out.dtype != np.uint8: | |
| raise ValueError(f"out must be uint8 {shape}, got {out.dtype} {out.shape}") | |
| out[r_h:] = g.pad_value | |
| out[:r_h, r_w:] = g.pad_value | |
| if r_h == 0 or r_w == 0: | |
| return out | |
| tw = self.col_w.astype(np.float64)[None, :, None] # exact float32 values | |
| cols0, cols1 = self.cols, self.cols + 1 | |
| for start in range(0, r_h, max(1, int(chunk_rows))): | |
| stop = min(r_h, start + int(chunk_rows)) | |
| hi = self.rows[start:stop] | |
| # horizontal pass on the source rows this chunk needs (hi and hi + 1), each row once | |
| need = np.unique(np.concatenate([hi, hi + 1])) | |
| src = img[need] # (R, src_w, C) uint8 | |
| a = src[:, cols0].astype(np.float64) # (R, r_w, C) | |
| b = src[:, cols1].astype(np.float64) | |
| inner = (a - tw * a).astype(F32) # fmaf(-t, a, a) | |
| horiz = (tw * b + inner).astype(F32) # fmaf(t, b, inner) | |
| horiz = np.trunc(horiz).astype(np.float64) # lerp1d(int a, int b, ...): truncation | |
| pos0 = np.searchsorted(need, hi) | |
| pos1 = np.searchsorted(need, hi + 1) | |
| a1, b1 = horiz[pos0], horiz[pos1] # (rows, r_w, C) | |
| th = self.row_w[start:stop].astype(np.float64)[:, None, None] | |
| inner = (a1 - th * a1).astype(F32) | |
| r = (th * b1 + inner).astype(F32) # second lerp1d, float32 | |
| out[start:stop, :r_w] = np.floor(r.astype(np.float64) + 0.5) # lroundf (r >= 0) | |
| return out | |
| _YOLOX_LUTS = LUTCache(maxsize=8) | |
| def yolox_letterbox_lut(src_h: int, src_w: int, dst_h: int = 960, dst_w: int = 960, | |
| pad_value: int = YOLOX_PAD_VALUE) -> YoloxLetterboxLUT: | |
| """The (cached) tables for one source size.""" | |
| def make() -> YoloxLetterboxLUT: | |
| g = yolox_letterbox_geometry(src_h, src_w, dst_h, dst_w, pad_value) | |
| k = F32(1.0 / np.float64(g.scale_f32)) # 1.0 / scale in double, passed as a float argument | |
| rows, row_w = _axis_table(dst_h, g.src_hw[0], k) | |
| cols, col_w = _axis_table(dst_w, g.src_hw[1], k) | |
| r_h, r_w = g.resized_hw | |
| return YoloxLetterboxLUT(g, rows[:r_h].copy(), row_w[:r_h].copy(), cols[:r_w].copy(), col_w[:r_w].copy(), | |
| float(k)) | |
| return _YOLOX_LUTS.get(("yolox", int(src_h), int(src_w), int(dst_h), int(dst_w), int(pad_value)), make) | |
| def yolox_letterbox(image: np.ndarray, dst_hw: Tuple[int, int] = (960, 960), *, pad_value: int = YOLOX_PAD_VALUE, | |
| layout: str = "hwc", out: Optional[np.ndarray] = None) -> Tuple[np.ndarray, LetterboxGeometry]: | |
| """Autoware YOLOX letterbox of one BGR8 image ``(H, W, 3)`` -> ``(uint8 image, geometry)``. | |
| ``layout``: ``"hwc"`` -> ``(dst_h, dst_w, 3)``; ``"nchw"`` -> ``(1, 3, dst_h, dst_w)`` (the ONNX ``images`` | |
| layout; ``.astype(np.float32)`` is exactly the network input). The channel order is unchanged (Autoware feeds | |
| BGR).""" | |
| img = np.asarray(image) | |
| lut = yolox_letterbox_lut(img.shape[0], img.shape[1], int(dst_hw[0]), int(dst_hw[1]), pad_value) | |
| if layout == "hwc": | |
| return lut.apply(img, out=out), lut.geometry | |
| if layout == "nchw": | |
| hwc = lut.apply(img) | |
| res = np.ascontiguousarray(hwc.transpose(2, 0, 1)[None]) | |
| if out is not None: | |
| out[...] = res | |
| return out, lut.geometry | |
| return res, lut.geometry | |
| raise ValueError(f"layout must be 'hwc' or 'nchw', got {layout!r}") | |
| # ------------------------------------------------------------------------------------------- BEVDet nearest crop | |
| # Per-plane statistics of the BEVDet Preprocess plugin, applied to the B, G, R planes in this order: training loaded | |
| # images as RGB (PIL) and its mmlabNormalize swapped them to BGR before applying the "RGB" ImageNet statistics, and | |
| # the Autoware node feeds BGR planes with the same numbers (research/bevdet/SPEC.md section 3.1). | |
| BEVDET_MEAN = (123.675, 116.28, 103.53) | |
| BEVDET_STD = (58.395, 57.12, 57.375) | |
| def _roundf(x: Any) -> np.ndarray: | |
| """C ``roundf`` (half away from zero) of float32 values -> int64.""" | |
| x = np.asarray(x, F32) | |
| return (np.sign(x) * np.floor(np.abs(x) + F32(0.5))).astype(np.int64) | |
| class NearestCropLUT: | |
| """The ``bevdet::Preprocess`` gather: output row ``i`` / column ``j`` reads source row ``rows[i]`` / column | |
| ``cols[j]`` of the (already 1600x900) image. Build with :func:`bevdet_nearest_lut` (cached); apply with | |
| :meth:`apply` (one HWC image) or :meth:`apply_planes` (planar ``[..., C, H, W]``).""" | |
| src_hw: Tuple[int, int] | |
| dst_hw: Tuple[int, int] | |
| crop_hw: Tuple[int, int] | |
| resize: float # exact float32 ratio dst_w / src_w, stored as a Python float | |
| rows: np.ndarray # (dst_h,) int64 | |
| cols: np.ndarray # (dst_w,) int64 | |
| def apply(self, image: np.ndarray, *, channels: str = "rgb", out: Optional[np.ndarray] = None) -> np.ndarray: | |
| """One ``(src_h, src_w, 3)`` uint8 image -> ``(3, dst_h, dst_w)`` uint8 B, G, R planes. ``channels`` is the | |
| input's order: ``"rgb"`` (PIL / ``ttaw.io.load_image``) or ``"bgr"`` (OpenCV, cv_bridge ``bgr8``).""" | |
| img = np.asarray(image) | |
| if img.ndim != 3 or img.shape[2] != 3 or img.dtype != np.uint8 or tuple(img.shape[:2]) != self.src_hw: | |
| raise ValueError(f"expected a uint8 ({self.src_hw[0]}, {self.src_hw[1]}, 3) image, got {img.dtype} " | |
| f"{img.shape} (resize it to the source size first, like the node's cv::resize)") | |
| if channels not in ("rgb", "bgr"): | |
| raise ValueError(f"channels must be 'rgb' or 'bgr', got {channels!r}") | |
| sub = img[self.rows][:, self.cols] # (dst_h, dst_w, 3) | |
| order = [2, 1, 0] if channels == "rgb" else [0, 1, 2] # -> B, G, R | |
| res = sub[:, :, order].transpose(2, 0, 1) | |
| if out is None: | |
| return np.ascontiguousarray(res) | |
| shape = (3,) + tuple(self.dst_hw) | |
| if out.shape != shape or out.dtype != np.uint8: | |
| raise ValueError(f"out must be uint8 {shape}, got {out.dtype} {out.shape}") | |
| out[...] = res | |
| return out | |
| def apply_planes(self, planes: np.ndarray) -> np.ndarray: | |
| """Planar ``[..., C, src_h, src_w]`` (any dtype) -> ``[..., C, dst_h, dst_w]``.""" | |
| p = np.asarray(planes) | |
| if tuple(p.shape[-2:]) != self.src_hw: | |
| raise ValueError(f"expected planes [..., {self.src_hw[0]}, {self.src_hw[1]}], got {p.shape}") | |
| return np.ascontiguousarray(p[..., self.rows, :][..., self.cols]) | |
| _BEVDET_LUTS = LUTCache(maxsize=4) | |
| def bevdet_nearest_lut(src_hw: Tuple[int, int] = (900, 1600), dst_hw: Tuple[int, int] = (256, 704), | |
| crop_hw: Tuple[int, int] = (140, 0)) -> NearestCropLUT: | |
| """The (cached) gather of the BEVDet Preprocess plugin: ``r = (float)dst_w / src_w`` (0.44f for the deployed | |
| config, also the ONNX attribute ``resize_radio``), ``rows[i] = roundf(i / r + crop_h / r)``, | |
| ``cols[j] = roundf(j / r + crop_w / r)``, every quantity float32 (deployed: rows 318, 320, 323, ..., 898 and | |
| columns 0, 2, 5, 7, ..., 1598). ``crop_hw`` is the crop offset in RESIZED pixels (the yaml ``crop``, (h, w)).""" | |
| src_h, src_w = int(src_hw[0]), int(src_hw[1]) | |
| dst_h, dst_w = int(dst_hw[0]), int(dst_hw[1]) | |
| crop_h, crop_w = int(crop_hw[0]), int(crop_hw[1]) | |
| def make() -> NearestCropLUT: | |
| r = F32(F32(dst_w) / F32(src_w)) | |
| off_h = F32(F32(crop_h) / r) | |
| off_w = F32(F32(crop_w) / r) | |
| rows = _roundf((np.arange(dst_h, dtype=F32) / r + off_h).astype(F32)) | |
| cols = _roundf((np.arange(dst_w, dtype=F32) / r + off_w).astype(F32)) | |
| if rows.min() < 0 or rows.max() >= src_h or cols.min() < 0 or cols.max() >= src_w: | |
| raise ValueError(f"crop {crop_hw} with ratio {float(r)} reads outside the {src_w}x{src_h} source") | |
| return NearestCropLUT((src_h, src_w), (dst_h, dst_w), (crop_h, crop_w), float(r), rows, cols) | |
| return _BEVDET_LUTS.get(("bevdet", src_h, src_w, dst_h, dst_w, crop_h, crop_w), make) | |
| def bevdet_nearest_crop(image: np.ndarray, *, channels: str = "rgb", src_hw: Tuple[int, int] = (900, 1600), | |
| dst_hw: Tuple[int, int] = (256, 704), crop_hw: Tuple[int, int] = (140, 0), | |
| out: Optional[np.ndarray] = None) -> np.ndarray: | |
| """One camera image at the source size (1600x900) -> its ``(3, 256, 704)`` uint8 B, G, R crop: exactly the | |
| samples the BEVDet network sees before ``bevdet_normalize``.""" | |
| return bevdet_nearest_lut(src_hw, dst_hw, crop_hw).apply(image, channels=channels, out=out) | |
| def bevdet_normalize(crop: np.ndarray, mean: Any = BEVDET_MEAN, std: Any = BEVDET_STD) -> np.ndarray: | |
| """``(x - mean[c]) / std[c]`` in float32 on B, G, R planes ``[..., 3, H, W]`` (the plugin's fp32 path; its fp16 | |
| mode computes in ``__half``).""" | |
| m = np.asarray(mean, F32).reshape(3, 1, 1) | |
| s = np.asarray(std, F32).reshape(3, 1, 1) | |
| return ((np.asarray(crop).astype(F32) - m) / s).astype(F32) | |
| # ------------------------------------------------------------------------------------------- BEVFormer | |
| # autoware_tensorrt_bevformer config/bevformer.param.yaml data_params (= the BEVFormer training img_norm_cfg, | |
| # DerryHub configs/bevformer/bevformer_small.py:19): caffe-style BGR means, std 1, to_rgb false | |
| BEVFORMER_MEAN_BGR = (103.530, 116.280, 123.675) | |
| BEVFORMER_STD = (1.0, 1.0, 1.0) | |
| class LinearResizeLUT: | |
| """OpenCV ``cv::resize(..., INTER_LINEAR)`` of float32 data for one (source, destination) size pair: row / column | |
| pairs and the float32 weight of the second tap (``resize.cpp`` coefficient loop; module docstring). Build with | |
| :func:`opencv_linear_lut` (cached); apply with :meth:`apply` (HWC) or :meth:`apply_planes` (``[..., H, W]``).""" | |
| src_hw: Tuple[int, int] | |
| dst_hw: Tuple[int, int] | |
| rows0: np.ndarray # (dst_h,) int64 | |
| rows1: np.ndarray # (dst_h,) int64 (clamped to the last row) | |
| wy: np.ndarray # (dst_h,) float32 weight of rows1 | |
| cols0: np.ndarray # (dst_w,) int64 | |
| cols1: np.ndarray | |
| wx: np.ndarray # (dst_w,) float32 weight of cols1 | |
| def _lerp(x0: np.ndarray, x1: np.ndarray, w: Any, out: Optional[np.ndarray] = None) -> np.ndarray: | |
| """``fma(x1 - x0, w, x0)`` in float32: the difference rounded once, the fused product-sum rounded once | |
| (float64 holds ``d * w + x0`` exactly for weights that are multiples of 1/8). ``x1`` is overwritten.""" | |
| d = np.subtract(x1, x0, out=x1) # float32: one rounding | |
| t = d.astype(np.float64) | |
| t *= w | |
| t += x0 | |
| if out is None: | |
| return t.astype(F32) | |
| out[...] = t # float64 -> float32: one rounding | |
| return out | |
| def _period(i0: np.ndarray, i1: np.ndarray, w: np.ndarray, n_src: int) -> Optional[Tuple[int, int]]: | |
| """``(p_in, p_out)`` when the taps repeat every ``p_out`` outputs shifted by ``p_in`` inputs, both taps of | |
| an output lie in one input block and no border clamp occurs (x0.8: 5 inputs -> 4 outputs), else None.""" | |
| n_dst = len(i0) | |
| for p_out in range(1, min(n_dst, 64) + 1): | |
| if n_dst % p_out or (n_src * p_out) % n_dst: | |
| continue | |
| p_in = n_src * p_out // n_dst | |
| base = (np.arange(n_dst) // p_out) * p_in | |
| reps = n_dst // p_out | |
| if (np.array_equal(i0 - base, np.tile(i0[:p_out], reps)) and np.array_equal(i1, i0 + 1) | |
| and np.array_equal(w, np.tile(w[:p_out], reps)) and int(i1[:p_out].max()) < p_in): | |
| return p_in, p_out | |
| return None | |
| def _pass(self, p: np.ndarray, axis: int, i0: np.ndarray, i1: np.ndarray, w: np.ndarray) -> np.ndarray: | |
| """One interpolation pass along ``axis`` (-1: columns, -2: rows) of a contiguous ``p`` [..., H, W]: strided | |
| block slices when the taps are periodic (no transposes), else gathers.""" | |
| period = self._period(i0, i1, w, p.shape[axis]) | |
| if period is None: | |
| wb = w.astype(np.float64) if axis == -1 else w.astype(np.float64)[:, None] | |
| return self._lerp(np.take(p, i0, axis=axis), np.take(p, i1, axis=axis), wb) | |
| p_in, p_out = period | |
| nb = p.shape[axis] // p_in | |
| taps0 = [int(t) for t in i0[:p_out]] | |
| run = taps0 == list(range(taps0[0], taps0[0] + p_out)) # x0.8: taps (0..3) and (1..4) of 5 | |
| wj = np.asarray(w[:p_out], np.float64) | |
| if axis == -1: | |
| blk = p.reshape(p.shape[:-1] + (nb, p_in)) # [..., H, nb, p_in] (a view) | |
| if run: | |
| a = taps0[0] | |
| res = self._lerp(blk[..., a:a + p_out], np.array(blk[..., a + 1:a + 1 + p_out]), wj) | |
| else: | |
| res = np.empty(p.shape[:-1] + (nb, p_out), F32) | |
| for j in range(p_out): | |
| self._lerp(blk[..., taps0[j]], np.array(blk[..., taps0[j] + 1]), wj[j], out=res[..., j]) | |
| return res.reshape(p.shape[:-1] + (nb * p_out,)) | |
| blk = p.reshape(p.shape[:-2] + (nb, p_in, p.shape[-1])) # [..., nb, p_in, W] (a view) | |
| if run: | |
| a = taps0[0] | |
| res = self._lerp(blk[..., a:a + p_out, :], np.array(blk[..., a + 1:a + 1 + p_out, :]), wj[:, None]) | |
| else: | |
| res = np.empty(p.shape[:-2] + (nb, p_out, p.shape[-1]), F32) | |
| for j in range(p_out): | |
| self._lerp(blk[..., taps0[j], :], np.array(blk[..., taps0[j] + 1, :]), wj[j], out=res[..., j, :]) | |
| return res.reshape(p.shape[:-2] + (nb * p_out, p.shape[-1])) | |
| def apply_planes(self, planes: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray: | |
| """Planar float32 ``[..., src_h, src_w]`` -> ``[..., dst_h, dst_w]`` (horizontal pass first, as OpenCV).""" | |
| p = np.ascontiguousarray(planes) | |
| if p.dtype != F32 or tuple(p.shape[-2:]) != self.src_hw: | |
| raise ValueError(f"expected float32 planes [..., {self.src_hw[0]}, {self.src_hw[1]}], got {p.dtype} " | |
| f"{p.shape}") | |
| shape = p.shape[:-2] + tuple(self.dst_hw) | |
| if out is not None and (out.shape != shape or out.dtype != F32): | |
| raise ValueError(f"out must be float32 {shape}, got {out.dtype} {out.shape}") | |
| h = self._pass(p, -1, self.cols0, self.cols1, self.wx) # [..., src_h, dst_w] | |
| v = self._pass(h, -2, self.rows0, self.rows1, self.wy) # [..., dst_h, dst_w] | |
| if out is None: | |
| return v | |
| out[...] = v | |
| return out | |
| def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray: | |
| """``(src_h, src_w[, C])`` float32 -> ``(dst_h, dst_w[, C])`` float32 (horizontal pass first, as OpenCV).""" | |
| img = np.asarray(image) | |
| if img.dtype != F32 or tuple(img.shape[:2]) != self.src_hw: | |
| raise ValueError(f"expected a float32 {self.src_hw} image, got {img.dtype} {img.shape}") | |
| planes = np.moveaxis(img, (0, 1), (-2, -1)) # [C..., H, W] | |
| res = np.moveaxis(self.apply_planes(planes), (-2, -1), (0, 1)) | |
| if out is None: | |
| return np.ascontiguousarray(res) | |
| if out.shape != res.shape or out.dtype != F32: | |
| raise ValueError(f"out must be float32 {res.shape}, got {out.dtype} {out.shape}") | |
| out[...] = res | |
| return out | |
| def _linear_axis(n_src: int, n_dst: int) -> Tuple[np.ndarray, np.ndarray, np.ndarray]: | |
| """``resize.cpp`` INTER_LINEAR coefficients of one axis: (first tap, second tap, float32 weight of the second).""" | |
| scale = 1.0 / (float(n_dst) / float(n_src)) # inv_scale = (double)dst / src; 1. / inv_scale | |
| f = ((np.arange(n_dst, dtype=np.float64) + 0.5) * scale - 0.5).astype(F32) | |
| s = np.floor(f).astype(np.int64) | |
| w = (f - s.astype(F32)).astype(F32) | |
| low = s < 0 | |
| s[low], w[low] = 0, F32(0) | |
| high = s >= n_src - 1 | |
| s[high], w[high] = n_src - 1, F32(0) | |
| return s, np.minimum(s + 1, n_src - 1), w | |
| def _eighths(w: np.ndarray) -> bool: | |
| x = np.asarray(w, np.float64) * 8.0 | |
| return bool(np.all(x == np.round(x))) | |
| _LINEAR_LUTS = LUTCache(maxsize=8) | |
| def opencv_linear_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int], *, strict: bool = True) -> LinearResizeLUT: | |
| """The (cached) float INTER_LINEAR tables for one size pair. | |
| Verified bit-exact against OpenCV 4.8.1 / 4.11.0 for every x0.8 down-scale (1600x900 -> 1280x720, the deployed | |
| one; 1920x1080, 160x90, 80x45), whose weights are multiples of 1/8. OpenCV computes the coefficients of other | |
| scales differently (up-scales and non-dyadic down-scales differ by up to ~1e-4 relative), so ``strict=True`` | |
| refuses a geometry whose weights are not multiples of 1/8; ``strict=False`` gives the textbook coefficients for | |
| those (an approximation, not an emulation). Exact 2x down-scales are always refused: OpenCV switches | |
| INTER_LINEAR to INTER_AREA there (``resize.cpp``).""" | |
| src_h, src_w = int(src_hw[0]), int(src_hw[1]) | |
| dst_h, dst_w = int(dst_hw[0]), int(dst_hw[1]) | |
| if min(src_h, src_w, dst_h, dst_w) < 1: | |
| raise ValueError(f"empty resize {src_hw} -> {dst_hw}") | |
| if src_h == 2 * dst_h and src_w == 2 * dst_w: | |
| raise ValueError("an exact 2x down-scale runs OpenCV's INTER_AREA path, which is not emulated") | |
| def make() -> LinearResizeLUT: | |
| r0, r1, wy = _linear_axis(src_h, dst_h) | |
| c0, c1, wx = _linear_axis(src_w, dst_w) | |
| return LinearResizeLUT((src_h, src_w), (dst_h, dst_w), r0, r1, wy, c0, c1, wx) | |
| lut = _LINEAR_LUTS.get(("linear_f32", src_h, src_w, dst_h, dst_w), make) | |
| if strict and not (_eighths(lut.wx) and _eighths(lut.wy)): | |
| raise ValueError(f"{src_hw} -> {dst_hw} is not a geometry this emulation reproduces bit for bit (weights must " | |
| "be multiples of 1/8, e.g. a x0.8 down-scale); pass strict=False for an approximation") | |
| return lut | |
| def opencv_resize_linear_f32(image: np.ndarray, dst_hw: Tuple[int, int], *, strict: bool = True) -> np.ndarray: | |
| """``cv2.resize(image, (dst_w, dst_h))`` (INTER_LINEAR) of a float32 ``(H, W[, C])`` image: bit-exact for the | |
| geometries :func:`opencv_linear_lut` accepts with ``strict=True``.""" | |
| img = np.asarray(image) | |
| return opencv_linear_lut(img.shape[:2], dst_hw, strict=strict).apply(img) | |
| def bevformer_input_geometry(src_hw: Tuple[int, int] = (900, 1600), scale: float = 0.8, | |
| pad_divisor: int = 32) -> Dict[str, Tuple[int, int]]: | |
| """Sizes of the node's pipeline for one source size: ``resized_hw`` = ``(int)(src * scale)`` in float32 | |
| (``augmentation_transforms.cpp:104-105``) and ``padded_hw`` (``PadMultiViewImages``, size divisor). Deployed: | |
| (900, 1600) -> (720, 1280) -> (736, 1280); the network normalises its UV by the padded size.""" | |
| s = F32(scale) | |
| h = int(F32(F32(int(src_hw[0])) * s)) | |
| w = int(F32(F32(int(src_hw[1])) * s)) | |
| d = int(pad_divisor) | |
| if d < 1: | |
| raise ValueError("pad_divisor must be >= 1") | |
| return {"src_hw": (int(src_hw[0]), int(src_hw[1])), "resized_hw": (h, w), | |
| "padded_hw": ((h + d - 1) // d * d, (w + d - 1) // d * d)} | |
| def _normalize_planes(image: np.ndarray, mean: Any, std: Any) -> np.ndarray: | |
| """BGR uint8 ``(H, W, 3)`` -> float32 planes ``(3, H, W)`` of ``(x - mean_c) / std_c`` as OpenCV evaluates it: | |
| ``convertTo`` with float ``alpha = 1/std``, ``beta = -mean/std`` and one fused multiply-add (exact in float64: | |
| an 8-bit value times a 24-bit alpha plus beta); for alpha == 1 that is the float addition ``x + beta``.""" | |
| img = np.asarray(image) | |
| if img.ndim != 3 or img.shape[2] != 3 or img.dtype != np.uint8: | |
| raise ValueError(f"expected a uint8 (H, W, 3) BGR image, got {img.dtype} {img.shape}") | |
| m = np.asarray(mean, np.float64).reshape(3) | |
| sd = np.asarray(std, np.float64).reshape(3) | |
| alpha = (1.0 / sd).astype(F32) | |
| beta = (-m / sd).astype(F32) | |
| planes = img.transpose(2, 0, 1).astype(F32) # exact: integers 0..255 | |
| for c in range(3): | |
| if alpha[c] == 1.0: | |
| planes[c] += beta[c] # float32, one rounding | |
| else: | |
| planes[c] = planes[c].astype(np.float64) * float(alpha[c]) + float(beta[c]) | |
| return planes | |
| def bevformer_normalize(image: np.ndarray, mean: Any = BEVFORMER_MEAN_BGR, std: Any = BEVFORMER_STD) -> np.ndarray: | |
| """One BGR uint8 ``(H, W, 3)`` image -> ``(x - mean_c) / std_c`` as float32 ``(H, W, 3)``, as OpenCV evaluates it | |
| (``convertTo`` with float ``alpha = 1/std``, ``beta = -mean/std``, one fused multiply-add; std 1 deployed).""" | |
| return np.ascontiguousarray(_normalize_planes(image, mean, std).transpose(1, 2, 0)) | |
| def bevformer_preprocess(images: Any, *, mean: Any = BEVFORMER_MEAN_BGR, std: Any = BEVFORMER_STD, | |
| scale: float = 0.8, pad_divisor: int = 32, out: Optional[np.ndarray] = None, | |
| workers: int = 1) -> np.ndarray: | |
| """N BGR uint8 camera images (a sequence of ``(H, W, 3)`` or an array ``[N, H, W, 3]``, one source size; the node | |
| feeds 1600x900 after its own ``cv::resize``) -> the network input ``[N, 3, Hp, Wp]`` float32: normalise, resize | |
| x``scale`` (OpenCV float INTER_LINEAR), zero-pad bottom / right to ``pad_divisor``, CHW (B, G, R planes, as | |
| ``to_rgb`` is false). Deployed: ``[6, 3, 736, 1280]`` (the ONNX input ``image`` is this with a leading 1). | |
| ``workers`` > 1 processes the cameras in that many threads (numpy releases the GIL in its loops; about 70 ms per | |
| 1600x900 camera single-threaded on the shared host).""" | |
| imgs = [np.asarray(im) for im in images] | |
| if not imgs: | |
| raise ValueError("no images") | |
| hw = tuple(imgs[0].shape[:2]) | |
| if any(tuple(im.shape[:2]) != hw for im in imgs): | |
| raise ValueError("every camera image must have the same size (the node resizes them to 1600x900 first)") | |
| geo = bevformer_input_geometry(hw, scale, pad_divisor) | |
| (rh, rw), (ph, pw) = geo["resized_hw"], geo["padded_hw"] | |
| lut = opencv_linear_lut(hw, (rh, rw)) # strict: the verified x0.8 geometries | |
| shape = (len(imgs), 3, ph, pw) | |
| if out is None: | |
| out = np.zeros(shape, F32) | |
| else: | |
| if out.shape != shape or out.dtype != F32: | |
| raise ValueError(f"out must be float32 {shape}, got {out.dtype} {out.shape}") | |
| out[:, :, rh:, :] = 0.0 | |
| out[:, :, :, rw:] = 0.0 | |
| def one(i: int) -> None: | |
| lut.apply_planes(_normalize_planes(imgs[i], mean, std), out=out[i, :, :rh, :rw]) | |
| if int(workers) > 1 and len(imgs) > 1: | |
| from concurrent.futures import ThreadPoolExecutor | |
| with ThreadPoolExecutor(max_workers=min(int(workers), len(imgs))) as pool: | |
| list(pool.map(one, range(len(imgs)))) | |
| else: | |
| for i in range(len(imgs)): | |
| one(i) | |
| return out | |
| # name -> (function, one-line description): the presets this module implements | |
| PRESETS: Dict[str, Tuple[Callable[..., Any], str]] = { | |
| "yolox_letterbox": (yolox_letterbox, "autoware_tensorrt_yolox preprocess.cu:43-129 letterbox (pad 114, BGR)"), | |
| "bevdet_nearest_crop": (bevdet_nearest_crop, "bevdet::Preprocess nearest resize 0.44 + crop (140, 0) of the " | |
| "1600x900 image -> B, G, R uint8 planes"), | |
| "bevformer_preprocess": (bevformer_preprocess, "autoware_tensorrt_bevformer normalise (BGR mean, std 1) -> " | |
| "cv::resize x0.8 INTER_LINEAR on float -> zero pad to /32, CHW"), | |
| } | |