changh95's picture
tt-model push diffusion-planner-p150 (container)
4d9b003 verified
Raw History Blame Contribute Delete
37.4 kB
# SPDX-License-Identifier: Apache-2.0
"""C14 image pre-processing: bit-exact host emulations of the Autoware camera nodes' resize kernels, driven by
per-source-size lookup tables (LUTs) that are computed once and cached.
Each preset reproduces one deployed CUDA / OpenCV routine exactly (integer outputs equal, not "close"), because a
standard resize in its place changes the network input enough to fail the PCC gates (YOLOX: ``cv2.resize`` drops the
detection-output PCC to 0.94, research/yolox/SPEC.md section 3). The tables hold everything that depends only on the
source and destination sizes (source indices, interpolation weights, letterbox geometry); applying them is a few
vectorised numpy passes over the image. numpy only; no ttnn, no torch, no OpenCV.
Presets (PLAN.md C14; the other camera ports add theirs here):
========================== =========================================================================================
``yolox_letterbox`` autoware_tensorrt_yolox ``resize_bilinear_letterbox_nhwc_to_nchw32_batch_kernel``
(autoware_universe @ 9ceaccf, ``perception/autoware_tensorrt_yolox/src/preprocess.cu:43-129``,
launched by ``tensorrt_yolox.cpp:243-332``): a "bilinear" resize with inverted,
half-magnitude weights, float -> int truncation between the horizontal and vertical passes,
``lroundf``, top-left letterbox padded with 114, BGR kept, no normalisation.
``bevdet_nearest_crop`` the ``bevdet::Preprocess`` TensorRT plugin of autoware_tensorrt_bevdet (``bevdet_vendor``
0.2.1; research/bevdet/SPEC.md section 3.2): nearest resize by r = 0.44f + crop (140, 0)
of the 1600x900 image, ``roundf(i / r + crop_h / r)`` rows / columns in float32, planar
B, G, R uint8 (the normalisation ``bevdet_normalize`` is separate: device-side in the TT
port). The node's ``cv::resize`` of non-1600x900 images to 1600x900 stays with the caller.
``bevformer_preprocess`` autoware_tensorrt_bevformer's ``preprocessing_pipeline`` (autoware_universe @ 9ceaccf,
``src/preprocessing/normalize_multiview_image.cpp:70-95`` +
``augmentation_transforms.cpp:90-222``; research/bevformer/SPEC.md section 3.1):
normalise FIRST (BGR uint8 -> float, ``(x - mean_c) / std_c``, BGR mean, std 1), then
``cv::resize`` x0.8 with INTER_LINEAR **on float data** (1600x900 -> 1280x720), then pad
bottom / right with 0 (normalised space) to a multiple of 32 (1280x736), CHW. The node's
``cv::resize`` of non-1600x900 uint8 images to 1600x900 stays with the caller.
========================== =========================================================================================
Exactness notes for ``yolox_letterbox`` (all arithmetic is float32 in the kernel):
- ``scale = min(dst_w / (float)src_w, dst_h / (float)src_h)``; ``r_h = (int)(scale * src_h)``,
``r_w = (int)(scale * src_w)``; the kernel receives ``k = (float)(1.0 / (double)scale)`` (``preprocess.cu:121-128``).
- per output row ``h``: ``c = k * (float)(h + 0.5)``, ``src_row = clamp(lroundf(c) - 1, 0, src_h - 2)``,
weight ``t = (1 + lroundf(c) - c) / 2`` (float32), and the same per column (``preprocess.cu:48-59, 77-88``).
- ``lerp1d(a, b, t) = fma(t, b, fma(-t, a, a))`` with float operands (CUDA's float overload, i.e. ``fmaf``:
one rounding per call); the horizontal results are truncated to int before the vertical ``lerp1d``
(``lerp1d`` takes ``int a, int b``), and the result is ``lroundf``-ed (``preprocess.cu:43-46, 57, 101-103``).
- ``t`` lies in [0.25, 0.75], so ``t`` is a multiple of 2**-25 and ``a - t*a`` / ``t*b + inner`` with integer
``a, b`` in [0, 255] need at most 34 significant bits: float64 holds them exactly, and one ``astype(float32)`` is
exactly the single rounding of ``fmaf``.
- Rows ``h >= r_h`` and columns ``w >= r_w`` are 114 (``preprocess.cu:109-110``); the output values are integers in
[0, 255], so this module returns uint8 (the network input is ``float(value)``, ``norm_factor`` 1.0).
Verified against two independent scalar transliterations of the CUDA source (``common/tests/host/test_image_host.py``
and research/yolox/scripts/{check_preprocess,verify_preproc_scalar}.py: 0 mismatches). Residual uncertainty, from the
SPEC: if nvcc picked the double ``fma`` overload, 0-2 of 3000 samples would differ by 1 (not testable without CUDA).
Exactness notes for ``bevformer_preprocess`` (float32 data; OpenCV's float INTER_LINEAR path):
- normalisation: ``Mat - mean`` / ``std`` evaluates as ``convertTo(alpha = 1/std, beta = -mean/std)`` with float
``alpha`` / ``beta`` (one fused multiply-add per value); with the deployed std 1 it is ``x - float(mean_c)``.
- resize coefficients (``resize.cpp``): ``scale = 1.0 / ((double)dst / src)``, ``f = (float)((d + 0.5) * scale -
0.5)``, ``s = floor(f)``, ``w = f - s`` (float), clamped to the border (``s < 0`` -> ``s = 0, w = 0``;
``s >= src - 1`` -> ``s = src - 1, w = 0``); the destination size is ``(int)(src * 0.8f)`` in float.
- each pass is ``fma(x1 - x0, w, x0)`` in float32 (the difference rounded once, then one rounding of the fused
product-sum), horizontal pass first, then vertical. This is what OpenCV 4.8.1 and 4.11.0 compute on this x86-64
host (AVX2 + FMA) bit for bit for every x0.8 down-scale, on arbitrary float images (not just uint8 - mean): the
plain two-product form ``x0 * (1 - w) + x1 * w`` differs from it in about 10 % of the samples by 1 ulp
(``tests/host/test_image_bevformer_host.py`` holds both checks). The weights of a x0.8 resize are multiples of
1/8, so float64 holds every fused product-sum exactly and one ``astype(float32)`` is the single rounding. OpenCV
computes the coefficients of other scales differently, so :func:`opencv_linear_lut` refuses them by default.
"""
from __future__ import annotations
import threading
from collections import OrderedDict
from dataclasses import dataclass
from typing import Any, Callable, Dict, Hashable, Optional, Tuple
import numpy as np
__all__ = [
"LetterboxGeometry",
"YoloxLetterboxLUT",
"yolox_letterbox",
"yolox_letterbox_geometry",
"yolox_letterbox_lut",
"lroundf",
"f32",
"LUTCache",
"PRESETS",
"YOLOX_PAD_VALUE",
"NearestCropLUT",
"bevdet_nearest_lut",
"bevdet_nearest_crop",
"bevdet_normalize",
"BEVDET_MEAN",
"BEVDET_STD",
"LinearResizeLUT",
"opencv_linear_lut",
"opencv_resize_linear_f32",
"bevformer_input_geometry",
"bevformer_normalize",
"bevformer_preprocess",
"BEVFORMER_MEAN_BGR",
"BEVFORMER_STD",
]
F32 = np.float32
YOLOX_PAD_VALUE = 114 # preprocess.cu:109-110 (letterbox fill), the YOLOX training pad value
def f32(x: Any) -> np.ndarray:
"""``x`` rounded once to float32 (round-to-nearest-even), as a numpy float32 scalar or array."""
return np.asarray(x).astype(F32)
def lroundf(x: Any) -> np.ndarray:
"""C ``lroundf``: round half away from zero, for float32 inputs (exact in float64). Returns int64."""
x = np.asarray(x, dtype=np.float64)
return (np.sign(x) * np.floor(np.abs(x) + 0.5)).astype(np.int64)
class LUTCache:
"""A small thread-safe LRU of tables keyed by the source / destination sizes (a camera stream keeps its size,
so the table is built once per stream; Autoware re-creates its buffers only when the source size changes,
``tensorrt_yolox.cpp:253-291``)."""
def __init__(self, maxsize: int = 8):
self.maxsize = int(maxsize)
self._items: "OrderedDict[Hashable, Any]" = OrderedDict()
self._lock = threading.Lock()
self.hits = 0
self.misses = 0
def get(self, key: Hashable, make: Callable[[], Any]) -> Any:
with self._lock:
if key in self._items:
self._items.move_to_end(key)
self.hits += 1
return self._items[key]
value = make() # outside the lock: building a table takes a few ms
with self._lock:
self.misses += 1
self._items[key] = value
self._items.move_to_end(key)
while len(self._items) > self.maxsize:
self._items.popitem(last=False)
return value
def clear(self) -> None:
with self._lock:
self._items.clear()
self.hits = self.misses = 0
def __len__(self) -> int:
return len(self._items)
# ------------------------------------------------------------------------------------------- YOLOX letterbox
@dataclass(frozen=True)
class LetterboxGeometry:
"""Where the source image lands in the network input, and what the decoder and the mask crop use.
``scale`` is the float32 scale of ``preprocess.cu:121`` (also ``scales_`` of ``tensorrt_yolox.cpp:285``, used to
map boxes back, and the mask scale of ``tensorrt_yolox.cpp:528-536``); ``resized_hw`` = ``(r_h, r_w)``, the
un-letterboxed region at the top-left of the ``dst_hw`` canvas. Autoware's mask crop ``(out_h, out_w)`` =
``((int)(src_h * scale), (int)(src_w * scale))`` is the same pair (float multiplication commutes)."""
src_hw: Tuple[int, int]
dst_hw: Tuple[int, int]
scale: float # exact float32 value, stored as a Python float
resized_hw: Tuple[int, int]
pad_value: int = YOLOX_PAD_VALUE
@property
def scale_f32(self) -> np.float32:
return F32(self.scale)
@property
def mask_hw(self) -> Tuple[int, int]:
"""``(out_h, out_w)`` of Autoware's mask (``getMaskImageGpu``): the un-letterboxed region."""
return self.resized_hw
def to_dict(self) -> Dict[str, Any]:
return {"src_hw": list(self.src_hw), "dst_hw": list(self.dst_hw), "scale": self.scale,
"scale_f32_hex": F32(self.scale).tobytes()[::-1].hex(), "resized_hw": list(self.resized_hw),
"mask_hw": list(self.mask_hw), "pad_value": self.pad_value,
"letterbox": "top-left (padding at the bottom / right)"}
def yolox_letterbox_geometry(src_h: int, src_w: int, dst_h: int = 960, dst_w: int = 960,
pad_value: int = YOLOX_PAD_VALUE) -> LetterboxGeometry:
"""Letterbox geometry of ``preprocess.cu:121-123`` (float32 scale, int truncation)."""
src_h, src_w, dst_h, dst_w = int(src_h), int(src_w), int(dst_h), int(dst_w)
if src_h < 2 or src_w < 2:
raise ValueError(f"the Autoware resize needs a source of at least 2x2 pixels, got {src_w}x{src_h}")
if dst_h < 1 or dst_w < 1:
raise ValueError(f"bad destination size {dst_w}x{dst_h}")
scale = min(F32(dst_w) / F32(src_w), F32(dst_h) / F32(src_h)) # float32 divisions
scale = F32(scale)
r_h = int(F32(scale * F32(src_h)))
r_w = int(F32(scale * F32(src_w)))
return LetterboxGeometry((src_h, src_w), (dst_h, dst_w), float(scale), (min(r_h, dst_h), min(r_w, dst_w)),
int(pad_value))
def _axis_table(n_dst: int, n_src: int, k: np.float32) -> Tuple[np.ndarray, np.ndarray]:
"""Source index and weight of every destination row (or column): ``preprocess.cu:48-51, 77-88``."""
c = (k * (np.arange(n_dst) + 0.5).astype(F32)).astype(F32) # centroid = scale * (float)(h + 0.5)
rc = lroundf(c)
idx = rc - 1
idx = np.where(idx < 0, 0, idx)
idx = np.where(idx >= n_src - 1, n_src - 2, idx)
weight = ((F32(1) + rc.astype(F32)).astype(F32) - c).astype(F32) # 1 + lroundf(c) - c (float32)
weight = (weight / F32(2)).astype(F32)
return idx.astype(np.int64), weight
@dataclass(frozen=True)
class YoloxLetterboxLUT:
"""The per-size tables of the YOLOX letterbox: source row / column of every destination row / column and the
float32 weights. Build with :func:`yolox_letterbox_lut` (cached); apply with :meth:`apply`."""
geometry: LetterboxGeometry
rows: np.ndarray # (r_h,) int64: source row hi (hi + 1 is the second tap)
row_w: np.ndarray # (r_h,) float32: vertical weight
cols: np.ndarray # (r_w,) int64: source column wi
col_w: np.ndarray # (r_w,) float32: horizontal weight
k: float # the kernel's float32 ``scale`` argument (1 / scale)
def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None, chunk_rows: int = 96) -> np.ndarray:
"""Letterbox one BGR uint8 image ``(src_h, src_w, 3)`` (any channel count) -> uint8 ``(dst_h, dst_w, C)``
HWC, channel order unchanged. ``out`` (uint8, that shape) is written in place when given."""
g = self.geometry
img = np.asarray(image)
if img.ndim != 3 or img.dtype != np.uint8 or tuple(img.shape[:2]) != g.src_hw:
raise ValueError(f"expected a uint8 ({g.src_hw[0]}, {g.src_hw[1]}, C) image, got {img.dtype} {img.shape}")
dst_h, dst_w = g.dst_hw
r_h, r_w = g.resized_hw
shape = (dst_h, dst_w, img.shape[2])
if out is None:
out = np.empty(shape, np.uint8)
elif out.shape != shape or out.dtype != np.uint8:
raise ValueError(f"out must be uint8 {shape}, got {out.dtype} {out.shape}")
out[r_h:] = g.pad_value
out[:r_h, r_w:] = g.pad_value
if r_h == 0 or r_w == 0:
return out
tw = self.col_w.astype(np.float64)[None, :, None] # exact float32 values
cols0, cols1 = self.cols, self.cols + 1
for start in range(0, r_h, max(1, int(chunk_rows))):
stop = min(r_h, start + int(chunk_rows))
hi = self.rows[start:stop]
# horizontal pass on the source rows this chunk needs (hi and hi + 1), each row once
need = np.unique(np.concatenate([hi, hi + 1]))
src = img[need] # (R, src_w, C) uint8
a = src[:, cols0].astype(np.float64) # (R, r_w, C)
b = src[:, cols1].astype(np.float64)
inner = (a - tw * a).astype(F32) # fmaf(-t, a, a)
horiz = (tw * b + inner).astype(F32) # fmaf(t, b, inner)
horiz = np.trunc(horiz).astype(np.float64) # lerp1d(int a, int b, ...): truncation
pos0 = np.searchsorted(need, hi)
pos1 = np.searchsorted(need, hi + 1)
a1, b1 = horiz[pos0], horiz[pos1] # (rows, r_w, C)
th = self.row_w[start:stop].astype(np.float64)[:, None, None]
inner = (a1 - th * a1).astype(F32)
r = (th * b1 + inner).astype(F32) # second lerp1d, float32
out[start:stop, :r_w] = np.floor(r.astype(np.float64) + 0.5) # lroundf (r >= 0)
return out
_YOLOX_LUTS = LUTCache(maxsize=8)
def yolox_letterbox_lut(src_h: int, src_w: int, dst_h: int = 960, dst_w: int = 960,
pad_value: int = YOLOX_PAD_VALUE) -> YoloxLetterboxLUT:
"""The (cached) tables for one source size."""
def make() -> YoloxLetterboxLUT:
g = yolox_letterbox_geometry(src_h, src_w, dst_h, dst_w, pad_value)
k = F32(1.0 / np.float64(g.scale_f32)) # 1.0 / scale in double, passed as a float argument
rows, row_w = _axis_table(dst_h, g.src_hw[0], k)
cols, col_w = _axis_table(dst_w, g.src_hw[1], k)
r_h, r_w = g.resized_hw
return YoloxLetterboxLUT(g, rows[:r_h].copy(), row_w[:r_h].copy(), cols[:r_w].copy(), col_w[:r_w].copy(),
float(k))
return _YOLOX_LUTS.get(("yolox", int(src_h), int(src_w), int(dst_h), int(dst_w), int(pad_value)), make)
def yolox_letterbox(image: np.ndarray, dst_hw: Tuple[int, int] = (960, 960), *, pad_value: int = YOLOX_PAD_VALUE,
layout: str = "hwc", out: Optional[np.ndarray] = None) -> Tuple[np.ndarray, LetterboxGeometry]:
"""Autoware YOLOX letterbox of one BGR8 image ``(H, W, 3)`` -> ``(uint8 image, geometry)``.
``layout``: ``"hwc"`` -> ``(dst_h, dst_w, 3)``; ``"nchw"`` -> ``(1, 3, dst_h, dst_w)`` (the ONNX ``images``
layout; ``.astype(np.float32)`` is exactly the network input). The channel order is unchanged (Autoware feeds
BGR)."""
img = np.asarray(image)
lut = yolox_letterbox_lut(img.shape[0], img.shape[1], int(dst_hw[0]), int(dst_hw[1]), pad_value)
if layout == "hwc":
return lut.apply(img, out=out), lut.geometry
if layout == "nchw":
hwc = lut.apply(img)
res = np.ascontiguousarray(hwc.transpose(2, 0, 1)[None])
if out is not None:
out[...] = res
return out, lut.geometry
return res, lut.geometry
raise ValueError(f"layout must be 'hwc' or 'nchw', got {layout!r}")
# ------------------------------------------------------------------------------------------- BEVDet nearest crop
# Per-plane statistics of the BEVDet Preprocess plugin, applied to the B, G, R planes in this order: training loaded
# images as RGB (PIL) and its mmlabNormalize swapped them to BGR before applying the "RGB" ImageNet statistics, and
# the Autoware node feeds BGR planes with the same numbers (research/bevdet/SPEC.md section 3.1).
BEVDET_MEAN = (123.675, 116.28, 103.53)
BEVDET_STD = (58.395, 57.12, 57.375)
def _roundf(x: Any) -> np.ndarray:
"""C ``roundf`` (half away from zero) of float32 values -> int64."""
x = np.asarray(x, F32)
return (np.sign(x) * np.floor(np.abs(x) + F32(0.5))).astype(np.int64)
@dataclass(frozen=True)
class NearestCropLUT:
"""The ``bevdet::Preprocess`` gather: output row ``i`` / column ``j`` reads source row ``rows[i]`` / column
``cols[j]`` of the (already 1600x900) image. Build with :func:`bevdet_nearest_lut` (cached); apply with
:meth:`apply` (one HWC image) or :meth:`apply_planes` (planar ``[..., C, H, W]``)."""
src_hw: Tuple[int, int]
dst_hw: Tuple[int, int]
crop_hw: Tuple[int, int]
resize: float # exact float32 ratio dst_w / src_w, stored as a Python float
rows: np.ndarray # (dst_h,) int64
cols: np.ndarray # (dst_w,) int64
def apply(self, image: np.ndarray, *, channels: str = "rgb", out: Optional[np.ndarray] = None) -> np.ndarray:
"""One ``(src_h, src_w, 3)`` uint8 image -> ``(3, dst_h, dst_w)`` uint8 B, G, R planes. ``channels`` is the
input's order: ``"rgb"`` (PIL / ``ttaw.io.load_image``) or ``"bgr"`` (OpenCV, cv_bridge ``bgr8``)."""
img = np.asarray(image)
if img.ndim != 3 or img.shape[2] != 3 or img.dtype != np.uint8 or tuple(img.shape[:2]) != self.src_hw:
raise ValueError(f"expected a uint8 ({self.src_hw[0]}, {self.src_hw[1]}, 3) image, got {img.dtype} "
f"{img.shape} (resize it to the source size first, like the node's cv::resize)")
if channels not in ("rgb", "bgr"):
raise ValueError(f"channels must be 'rgb' or 'bgr', got {channels!r}")
sub = img[self.rows][:, self.cols] # (dst_h, dst_w, 3)
order = [2, 1, 0] if channels == "rgb" else [0, 1, 2] # -> B, G, R
res = sub[:, :, order].transpose(2, 0, 1)
if out is None:
return np.ascontiguousarray(res)
shape = (3,) + tuple(self.dst_hw)
if out.shape != shape or out.dtype != np.uint8:
raise ValueError(f"out must be uint8 {shape}, got {out.dtype} {out.shape}")
out[...] = res
return out
def apply_planes(self, planes: np.ndarray) -> np.ndarray:
"""Planar ``[..., C, src_h, src_w]`` (any dtype) -> ``[..., C, dst_h, dst_w]``."""
p = np.asarray(planes)
if tuple(p.shape[-2:]) != self.src_hw:
raise ValueError(f"expected planes [..., {self.src_hw[0]}, {self.src_hw[1]}], got {p.shape}")
return np.ascontiguousarray(p[..., self.rows, :][..., self.cols])
_BEVDET_LUTS = LUTCache(maxsize=4)
def bevdet_nearest_lut(src_hw: Tuple[int, int] = (900, 1600), dst_hw: Tuple[int, int] = (256, 704),
crop_hw: Tuple[int, int] = (140, 0)) -> NearestCropLUT:
"""The (cached) gather of the BEVDet Preprocess plugin: ``r = (float)dst_w / src_w`` (0.44f for the deployed
config, also the ONNX attribute ``resize_radio``), ``rows[i] = roundf(i / r + crop_h / r)``,
``cols[j] = roundf(j / r + crop_w / r)``, every quantity float32 (deployed: rows 318, 320, 323, ..., 898 and
columns 0, 2, 5, 7, ..., 1598). ``crop_hw`` is the crop offset in RESIZED pixels (the yaml ``crop``, (h, w))."""
src_h, src_w = int(src_hw[0]), int(src_hw[1])
dst_h, dst_w = int(dst_hw[0]), int(dst_hw[1])
crop_h, crop_w = int(crop_hw[0]), int(crop_hw[1])
def make() -> NearestCropLUT:
r = F32(F32(dst_w) / F32(src_w))
off_h = F32(F32(crop_h) / r)
off_w = F32(F32(crop_w) / r)
rows = _roundf((np.arange(dst_h, dtype=F32) / r + off_h).astype(F32))
cols = _roundf((np.arange(dst_w, dtype=F32) / r + off_w).astype(F32))
if rows.min() < 0 or rows.max() >= src_h or cols.min() < 0 or cols.max() >= src_w:
raise ValueError(f"crop {crop_hw} with ratio {float(r)} reads outside the {src_w}x{src_h} source")
return NearestCropLUT((src_h, src_w), (dst_h, dst_w), (crop_h, crop_w), float(r), rows, cols)
return _BEVDET_LUTS.get(("bevdet", src_h, src_w, dst_h, dst_w, crop_h, crop_w), make)
def bevdet_nearest_crop(image: np.ndarray, *, channels: str = "rgb", src_hw: Tuple[int, int] = (900, 1600),
dst_hw: Tuple[int, int] = (256, 704), crop_hw: Tuple[int, int] = (140, 0),
out: Optional[np.ndarray] = None) -> np.ndarray:
"""One camera image at the source size (1600x900) -> its ``(3, 256, 704)`` uint8 B, G, R crop: exactly the
samples the BEVDet network sees before ``bevdet_normalize``."""
return bevdet_nearest_lut(src_hw, dst_hw, crop_hw).apply(image, channels=channels, out=out)
def bevdet_normalize(crop: np.ndarray, mean: Any = BEVDET_MEAN, std: Any = BEVDET_STD) -> np.ndarray:
"""``(x - mean[c]) / std[c]`` in float32 on B, G, R planes ``[..., 3, H, W]`` (the plugin's fp32 path; its fp16
mode computes in ``__half``)."""
m = np.asarray(mean, F32).reshape(3, 1, 1)
s = np.asarray(std, F32).reshape(3, 1, 1)
return ((np.asarray(crop).astype(F32) - m) / s).astype(F32)
# ------------------------------------------------------------------------------------------- BEVFormer
# autoware_tensorrt_bevformer config/bevformer.param.yaml data_params (= the BEVFormer training img_norm_cfg,
# DerryHub configs/bevformer/bevformer_small.py:19): caffe-style BGR means, std 1, to_rgb false
BEVFORMER_MEAN_BGR = (103.530, 116.280, 123.675)
BEVFORMER_STD = (1.0, 1.0, 1.0)
@dataclass(frozen=True)
class LinearResizeLUT:
"""OpenCV ``cv::resize(..., INTER_LINEAR)`` of float32 data for one (source, destination) size pair: row / column
pairs and the float32 weight of the second tap (``resize.cpp`` coefficient loop; module docstring). Build with
:func:`opencv_linear_lut` (cached); apply with :meth:`apply` (HWC) or :meth:`apply_planes` (``[..., H, W]``)."""
src_hw: Tuple[int, int]
dst_hw: Tuple[int, int]
rows0: np.ndarray # (dst_h,) int64
rows1: np.ndarray # (dst_h,) int64 (clamped to the last row)
wy: np.ndarray # (dst_h,) float32 weight of rows1
cols0: np.ndarray # (dst_w,) int64
cols1: np.ndarray
wx: np.ndarray # (dst_w,) float32 weight of cols1
@staticmethod
def _lerp(x0: np.ndarray, x1: np.ndarray, w: Any, out: Optional[np.ndarray] = None) -> np.ndarray:
"""``fma(x1 - x0, w, x0)`` in float32: the difference rounded once, the fused product-sum rounded once
(float64 holds ``d * w + x0`` exactly for weights that are multiples of 1/8). ``x1`` is overwritten."""
d = np.subtract(x1, x0, out=x1) # float32: one rounding
t = d.astype(np.float64)
t *= w
t += x0
if out is None:
return t.astype(F32)
out[...] = t # float64 -> float32: one rounding
return out
@staticmethod
def _period(i0: np.ndarray, i1: np.ndarray, w: np.ndarray, n_src: int) -> Optional[Tuple[int, int]]:
"""``(p_in, p_out)`` when the taps repeat every ``p_out`` outputs shifted by ``p_in`` inputs, both taps of
an output lie in one input block and no border clamp occurs (x0.8: 5 inputs -> 4 outputs), else None."""
n_dst = len(i0)
for p_out in range(1, min(n_dst, 64) + 1):
if n_dst % p_out or (n_src * p_out) % n_dst:
continue
p_in = n_src * p_out // n_dst
base = (np.arange(n_dst) // p_out) * p_in
reps = n_dst // p_out
if (np.array_equal(i0 - base, np.tile(i0[:p_out], reps)) and np.array_equal(i1, i0 + 1)
and np.array_equal(w, np.tile(w[:p_out], reps)) and int(i1[:p_out].max()) < p_in):
return p_in, p_out
return None
def _pass(self, p: np.ndarray, axis: int, i0: np.ndarray, i1: np.ndarray, w: np.ndarray) -> np.ndarray:
"""One interpolation pass along ``axis`` (-1: columns, -2: rows) of a contiguous ``p`` [..., H, W]: strided
block slices when the taps are periodic (no transposes), else gathers."""
period = self._period(i0, i1, w, p.shape[axis])
if period is None:
wb = w.astype(np.float64) if axis == -1 else w.astype(np.float64)[:, None]
return self._lerp(np.take(p, i0, axis=axis), np.take(p, i1, axis=axis), wb)
p_in, p_out = period
nb = p.shape[axis] // p_in
taps0 = [int(t) for t in i0[:p_out]]
run = taps0 == list(range(taps0[0], taps0[0] + p_out)) # x0.8: taps (0..3) and (1..4) of 5
wj = np.asarray(w[:p_out], np.float64)
if axis == -1:
blk = p.reshape(p.shape[:-1] + (nb, p_in)) # [..., H, nb, p_in] (a view)
if run:
a = taps0[0]
res = self._lerp(blk[..., a:a + p_out], np.array(blk[..., a + 1:a + 1 + p_out]), wj)
else:
res = np.empty(p.shape[:-1] + (nb, p_out), F32)
for j in range(p_out):
self._lerp(blk[..., taps0[j]], np.array(blk[..., taps0[j] + 1]), wj[j], out=res[..., j])
return res.reshape(p.shape[:-1] + (nb * p_out,))
blk = p.reshape(p.shape[:-2] + (nb, p_in, p.shape[-1])) # [..., nb, p_in, W] (a view)
if run:
a = taps0[0]
res = self._lerp(blk[..., a:a + p_out, :], np.array(blk[..., a + 1:a + 1 + p_out, :]), wj[:, None])
else:
res = np.empty(p.shape[:-2] + (nb, p_out, p.shape[-1]), F32)
for j in range(p_out):
self._lerp(blk[..., taps0[j], :], np.array(blk[..., taps0[j] + 1, :]), wj[j], out=res[..., j, :])
return res.reshape(p.shape[:-2] + (nb * p_out, p.shape[-1]))
def apply_planes(self, planes: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray:
"""Planar float32 ``[..., src_h, src_w]`` -> ``[..., dst_h, dst_w]`` (horizontal pass first, as OpenCV)."""
p = np.ascontiguousarray(planes)
if p.dtype != F32 or tuple(p.shape[-2:]) != self.src_hw:
raise ValueError(f"expected float32 planes [..., {self.src_hw[0]}, {self.src_hw[1]}], got {p.dtype} "
f"{p.shape}")
shape = p.shape[:-2] + tuple(self.dst_hw)
if out is not None and (out.shape != shape or out.dtype != F32):
raise ValueError(f"out must be float32 {shape}, got {out.dtype} {out.shape}")
h = self._pass(p, -1, self.cols0, self.cols1, self.wx) # [..., src_h, dst_w]
v = self._pass(h, -2, self.rows0, self.rows1, self.wy) # [..., dst_h, dst_w]
if out is None:
return v
out[...] = v
return out
def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray:
"""``(src_h, src_w[, C])`` float32 -> ``(dst_h, dst_w[, C])`` float32 (horizontal pass first, as OpenCV)."""
img = np.asarray(image)
if img.dtype != F32 or tuple(img.shape[:2]) != self.src_hw:
raise ValueError(f"expected a float32 {self.src_hw} image, got {img.dtype} {img.shape}")
planes = np.moveaxis(img, (0, 1), (-2, -1)) # [C..., H, W]
res = np.moveaxis(self.apply_planes(planes), (-2, -1), (0, 1))
if out is None:
return np.ascontiguousarray(res)
if out.shape != res.shape or out.dtype != F32:
raise ValueError(f"out must be float32 {res.shape}, got {out.dtype} {out.shape}")
out[...] = res
return out
def _linear_axis(n_src: int, n_dst: int) -> Tuple[np.ndarray, np.ndarray, np.ndarray]:
"""``resize.cpp`` INTER_LINEAR coefficients of one axis: (first tap, second tap, float32 weight of the second)."""
scale = 1.0 / (float(n_dst) / float(n_src)) # inv_scale = (double)dst / src; 1. / inv_scale
f = ((np.arange(n_dst, dtype=np.float64) + 0.5) * scale - 0.5).astype(F32)
s = np.floor(f).astype(np.int64)
w = (f - s.astype(F32)).astype(F32)
low = s < 0
s[low], w[low] = 0, F32(0)
high = s >= n_src - 1
s[high], w[high] = n_src - 1, F32(0)
return s, np.minimum(s + 1, n_src - 1), w
def _eighths(w: np.ndarray) -> bool:
x = np.asarray(w, np.float64) * 8.0
return bool(np.all(x == np.round(x)))
_LINEAR_LUTS = LUTCache(maxsize=8)
def opencv_linear_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int], *, strict: bool = True) -> LinearResizeLUT:
"""The (cached) float INTER_LINEAR tables for one size pair.
Verified bit-exact against OpenCV 4.8.1 / 4.11.0 for every x0.8 down-scale (1600x900 -> 1280x720, the deployed
one; 1920x1080, 160x90, 80x45), whose weights are multiples of 1/8. OpenCV computes the coefficients of other
scales differently (up-scales and non-dyadic down-scales differ by up to ~1e-4 relative), so ``strict=True``
refuses a geometry whose weights are not multiples of 1/8; ``strict=False`` gives the textbook coefficients for
those (an approximation, not an emulation). Exact 2x down-scales are always refused: OpenCV switches
INTER_LINEAR to INTER_AREA there (``resize.cpp``)."""
src_h, src_w = int(src_hw[0]), int(src_hw[1])
dst_h, dst_w = int(dst_hw[0]), int(dst_hw[1])
if min(src_h, src_w, dst_h, dst_w) < 1:
raise ValueError(f"empty resize {src_hw} -> {dst_hw}")
if src_h == 2 * dst_h and src_w == 2 * dst_w:
raise ValueError("an exact 2x down-scale runs OpenCV's INTER_AREA path, which is not emulated")
def make() -> LinearResizeLUT:
r0, r1, wy = _linear_axis(src_h, dst_h)
c0, c1, wx = _linear_axis(src_w, dst_w)
return LinearResizeLUT((src_h, src_w), (dst_h, dst_w), r0, r1, wy, c0, c1, wx)
lut = _LINEAR_LUTS.get(("linear_f32", src_h, src_w, dst_h, dst_w), make)
if strict and not (_eighths(lut.wx) and _eighths(lut.wy)):
raise ValueError(f"{src_hw} -> {dst_hw} is not a geometry this emulation reproduces bit for bit (weights must "
"be multiples of 1/8, e.g. a x0.8 down-scale); pass strict=False for an approximation")
return lut
def opencv_resize_linear_f32(image: np.ndarray, dst_hw: Tuple[int, int], *, strict: bool = True) -> np.ndarray:
"""``cv2.resize(image, (dst_w, dst_h))`` (INTER_LINEAR) of a float32 ``(H, W[, C])`` image: bit-exact for the
geometries :func:`opencv_linear_lut` accepts with ``strict=True``."""
img = np.asarray(image)
return opencv_linear_lut(img.shape[:2], dst_hw, strict=strict).apply(img)
def bevformer_input_geometry(src_hw: Tuple[int, int] = (900, 1600), scale: float = 0.8,
pad_divisor: int = 32) -> Dict[str, Tuple[int, int]]:
"""Sizes of the node's pipeline for one source size: ``resized_hw`` = ``(int)(src * scale)`` in float32
(``augmentation_transforms.cpp:104-105``) and ``padded_hw`` (``PadMultiViewImages``, size divisor). Deployed:
(900, 1600) -> (720, 1280) -> (736, 1280); the network normalises its UV by the padded size."""
s = F32(scale)
h = int(F32(F32(int(src_hw[0])) * s))
w = int(F32(F32(int(src_hw[1])) * s))
d = int(pad_divisor)
if d < 1:
raise ValueError("pad_divisor must be >= 1")
return {"src_hw": (int(src_hw[0]), int(src_hw[1])), "resized_hw": (h, w),
"padded_hw": ((h + d - 1) // d * d, (w + d - 1) // d * d)}
def _normalize_planes(image: np.ndarray, mean: Any, std: Any) -> np.ndarray:
"""BGR uint8 ``(H, W, 3)`` -> float32 planes ``(3, H, W)`` of ``(x - mean_c) / std_c`` as OpenCV evaluates it:
``convertTo`` with float ``alpha = 1/std``, ``beta = -mean/std`` and one fused multiply-add (exact in float64:
an 8-bit value times a 24-bit alpha plus beta); for alpha == 1 that is the float addition ``x + beta``."""
img = np.asarray(image)
if img.ndim != 3 or img.shape[2] != 3 or img.dtype != np.uint8:
raise ValueError(f"expected a uint8 (H, W, 3) BGR image, got {img.dtype} {img.shape}")
m = np.asarray(mean, np.float64).reshape(3)
sd = np.asarray(std, np.float64).reshape(3)
alpha = (1.0 / sd).astype(F32)
beta = (-m / sd).astype(F32)
planes = img.transpose(2, 0, 1).astype(F32) # exact: integers 0..255
for c in range(3):
if alpha[c] == 1.0:
planes[c] += beta[c] # float32, one rounding
else:
planes[c] = planes[c].astype(np.float64) * float(alpha[c]) + float(beta[c])
return planes
def bevformer_normalize(image: np.ndarray, mean: Any = BEVFORMER_MEAN_BGR, std: Any = BEVFORMER_STD) -> np.ndarray:
"""One BGR uint8 ``(H, W, 3)`` image -> ``(x - mean_c) / std_c`` as float32 ``(H, W, 3)``, as OpenCV evaluates it
(``convertTo`` with float ``alpha = 1/std``, ``beta = -mean/std``, one fused multiply-add; std 1 deployed)."""
return np.ascontiguousarray(_normalize_planes(image, mean, std).transpose(1, 2, 0))
def bevformer_preprocess(images: Any, *, mean: Any = BEVFORMER_MEAN_BGR, std: Any = BEVFORMER_STD,
scale: float = 0.8, pad_divisor: int = 32, out: Optional[np.ndarray] = None,
workers: int = 1) -> np.ndarray:
"""N BGR uint8 camera images (a sequence of ``(H, W, 3)`` or an array ``[N, H, W, 3]``, one source size; the node
feeds 1600x900 after its own ``cv::resize``) -> the network input ``[N, 3, Hp, Wp]`` float32: normalise, resize
x``scale`` (OpenCV float INTER_LINEAR), zero-pad bottom / right to ``pad_divisor``, CHW (B, G, R planes, as
``to_rgb`` is false). Deployed: ``[6, 3, 736, 1280]`` (the ONNX input ``image`` is this with a leading 1).
``workers`` > 1 processes the cameras in that many threads (numpy releases the GIL in its loops; about 70 ms per
1600x900 camera single-threaded on the shared host)."""
imgs = [np.asarray(im) for im in images]
if not imgs:
raise ValueError("no images")
hw = tuple(imgs[0].shape[:2])
if any(tuple(im.shape[:2]) != hw for im in imgs):
raise ValueError("every camera image must have the same size (the node resizes them to 1600x900 first)")
geo = bevformer_input_geometry(hw, scale, pad_divisor)
(rh, rw), (ph, pw) = geo["resized_hw"], geo["padded_hw"]
lut = opencv_linear_lut(hw, (rh, rw)) # strict: the verified x0.8 geometries
shape = (len(imgs), 3, ph, pw)
if out is None:
out = np.zeros(shape, F32)
else:
if out.shape != shape or out.dtype != F32:
raise ValueError(f"out must be float32 {shape}, got {out.dtype} {out.shape}")
out[:, :, rh:, :] = 0.0
out[:, :, :, rw:] = 0.0
def one(i: int) -> None:
lut.apply_planes(_normalize_planes(imgs[i], mean, std), out=out[i, :, :rh, :rw])
if int(workers) > 1 and len(imgs) > 1:
from concurrent.futures import ThreadPoolExecutor
with ThreadPoolExecutor(max_workers=min(int(workers), len(imgs))) as pool:
list(pool.map(one, range(len(imgs))))
else:
for i in range(len(imgs)):
one(i)
return out
# name -> (function, one-line description): the presets this module implements
PRESETS: Dict[str, Tuple[Callable[..., Any], str]] = {
"yolox_letterbox": (yolox_letterbox, "autoware_tensorrt_yolox preprocess.cu:43-129 letterbox (pad 114, BGR)"),
"bevdet_nearest_crop": (bevdet_nearest_crop, "bevdet::Preprocess nearest resize 0.44 + crop (140, 0) of the "
"1600x900 image -> B, G, R uint8 planes"),
"bevformer_preprocess": (bevformer_preprocess, "autoware_tensorrt_bevformer normalise (BGR mean, std 1) -> "
"cv::resize x0.8 INTER_LINEAR on float -> zero pad to /32, CHW"),
}