# tt_superpoint: Python API Use SuperPoint (keypoint detection and 256-d descriptors) on one Tenstorrent Blackhole chip from Python code, notebooks and pipelines. The HTTP server (`models/server/app.py`, see `SERVING.md`) runs the same request path. ## Install Install the package into a Python environment that has ttnn (tt-metal). Run the commands from the repo root: ```bash pip install -e code/ # the Python API only (or: pip install code/) pip install -e "code/[server,test]" # also the HTTP server (fastapi, uvicorn, pydantic) and pytest ``` The package declares only its own dependencies (numpy, torch, pillow, transformers, huggingface_hub, safetensors). ttnn is not a PyPI dependency. You do not need `PYTHONPATH` or `sys.path` changes. | to run | extras | needs | |---|---|---| | the Python API | none | ttnn and a chip | | the HTTP server (`models/server/app.py`, see `SERVING.md`) | `server` | ttnn and a chip | | the host tests (no chip) | `server,test` | ttnn importable (the tests compare with the server code and import the port modules), no device | | the device tests | `server,test` | ttnn and a chip | Run only the host tests (no device is opened): ```bash cd code && TT_VISIBLE_DEVICES=none python -m pytest -q models/tests/test_api_host.py models/tests/test_api_warmup_host.py ``` The device tests of the API are `models/tests/test_api_device.py` and `models/tests/test_api_warmup_device.py` (`cd code && python -m pytest -s -q `). The weights (`magic-leap-community/superpoint`, 5 MB) come from the Hugging Face cache. The first call downloads them if they are not in the cache. ## Quickstart ```python from tt_superpoint import SuperPoint with SuperPoint.from_pretrained(device_id=0) as model: out = model("code/sample_data/house_in_field_1080p.jpg") print(len(out), out.keypoints[:3], out.scores[:3], out.descriptors.shape) ``` The relative path `code/sample_data/...` in the snippets of this file assumes that the current directory is the repo root. From another directory, use an absolute path. `python code/examples/quickstart.py` runs this snippet on the demo image. The script finds the demo image from its own location, so you can run it from any directory. It writes `keypoints.png`, `keypoints.json` and `descriptors.npy` to `quickstart_out/` in the current directory (`--out-dir` changes this). ## `SuperPoint.from_pretrained(...)` ```python SuperPoint.from_pretrained( pretrained_model_name_or_path="magic-leap-community/superpoint", *, revision=None, device=None, device_id=0, dispatch="auto", warmup_variants=None, # None = "default"; see "Warm-up" below precompile_sizes=None, precompile_nms_radii=(), trace_region_size=32 << 20, local_files_only=False, verbose=False, ) -> SuperPoint ``` This call loads the weights, opens the chip, builds the model, compiles the kernels, captures the metal traces and warms up the per-request variants (see [Warm-up](#warm-up)). Startup takes approximately 5 s with cached kernels and 10-60 s without them. | argument | meaning | |---|---| | `pretrained_model_name_or_path` | Hugging Face repo id, or a local directory with `config.json` and `model.safetensors`. | | `revision` | Weights revision. The default for the default repo is `734450e9`, the revision of the published numbers. | | `device` | An open ttnn device. The device must have `l1_small_size >= 32768` and `trace_region_size >= 32 MiB`. `close()` does not close this device. | | `device_id` | The chip to open when `device` is `None`. | | `dispatch` | `"auto"` (default): ETH dispatch, 1 command queue and a 12x10 compute grid when the tt-metal tree has the ETH-dispatch patch. This is the p150 target configuration and the configuration of all published numbers. Without the patch, the default (Tensix) dispatch is used and a `RuntimeWarning` is shown; on a p150 the compute grid is then 11x10. The outputs are the same. `"eth"` forces ETH dispatch. `"worker"` forces the Tensix dispatch: an explicit opt-in for comparisons on a Galaxy chip (12x10 there, 11x10 on a p150). `$SP_DISPATCH` overrides `"auto"`. The model uses only command queue 0. | | `warmup_variants` | The per-request variants to prepare at startup, so that the first real call of each one is as fast as the later calls. See [Warm-up](#warm-up). | | `precompile_sizes` | Older name. When given, it replaces `warmup_variants["sizes"]`. | | `precompile_nms_radii` | Older name. The values are added to `warmup_variants["nms_radii"]`. | | `trace_region_size` | Trace region in bytes when this call opens the device. | | `local_files_only` | Load the weights from the Hugging Face cache only (no network). | | `verbose` | `False` (default): tt-metal and ttnn show only warnings and errors. `True`: they also show their info log. See [Console output](#console-output). | ## `model(images, **params)` ```python out = model(image, max_keypoints=1024, keypoint_threshold=0.005, nms_radius=4, return_descriptors=True, bgr=False) # -> SuperPointOutput outs = model([image1, image2, ...], num_workers=4) # -> list[SuperPointOutput] ``` Images: | input | notes | |---|---| | `str` / `pathlib.Path` | An image file (PNG, JPEG, or any format that Pillow opens). | | `bytes` | Encoded file contents. | | `PIL.Image.Image` | Any mode. | | `numpy.ndarray` | `(H, W)` gray, or `(H, W, C)` with C = 1, 3 or 4 (RGB or RGBA). Use `bgr=True` for OpenCV (`cv2.imread`) arrays. | | `torch.Tensor` | `(H, W)`, `(C, H, W)` (torchvision) or `(H, W, C)`. | | a list or tuple | One output for each image, in input order. | | 4-D array or tensor | A batch: `(B, H, W, C)` numpy or `(B, C, H, W)` torch. One output for each image. | - dtypes: uint8 (0..255), other integers in 0..255, or floats in [0, 1]. The model rounds float values to 8 bits, because the device reads 8-bit images. - Any image size is accepted. The model resizes every image to the 480x640 network frame (bilinear, the aspect ratio is not kept). The resize is bit-exact with Pillow. It runs on the device when the size fits the device resize kernel, for example 3840x2160 and the smaller common sizes (one resize trace for each size, built at startup for the warm-up sizes and on first use for other sizes, 8 sizes kept). Other sizes are resized on the host, for example 4032x3024 (12 MP phone photos). - The model reads channel 0 (R) of the RGB image, like the Hugging Face `SuperPointForKeypointDetection`. For a grayscale image, channel 0 is the gray plane. Parameters (the same defaults and limits as the HTTP server): | parameter | default | range | effect | |---|---|---|---| | `max_keypoints` | 1024 | -1..307200 | Keep the top-k keypoints by score. -1 keeps all keypoints above the threshold. | | `keypoint_threshold` | 0.005 | [0, 1] | Minimum score after NMS. | | `nms_radius` | 4 | 0..32 | NMS radius in network-frame pixels. Radii 1..8 run on the device (4 is in the main trace; other radii have a separate trace). Radius 0 (no NMS) and radii above 8 run on the host, which is slower. | | `return_descriptors` | True | bool | When False, `descriptors` is `None`. | | `bgr` | False | bool | Arrays are in BGR order (OpenCV). | | `num_workers` | 4 | int >= 0 | Lists only: host threads that decode the next images while the device runs the current image. 0 = no threads. | `model.iter(images, **params)` gives the same outputs as `model(list)` one at a time. It accepts any iterable, for example a generator over video frames. ## `SuperPointOutput` | field | type | content | |---|---|---| | `keypoints` | `torch.float32 (N, 2)` | `[x, y]` in pixels of the original image (network-frame position multiplied by `scale`). | | `scores` | `torch.float32 (N,)` | Keypoint scores in (0, 1], sorted in descending order. `keypoints` and `descriptors` use the same order. | | `descriptors` | `torch.float32 (N, 256)` or `None` | L2-normalised descriptors, bilinearly sampled at the keypoints. | | `image_size` | `(int, int)` | `(height, width)` of the original image. | | `scale` | `(float, float)` | `(sx, sy)` = original size / (640, 480). | | `device_nms` | `bool` | True when NMS ran on the device. | | `timing_ms` | `dict` | `preprocess`, `device_forward`, `postprocess` of this call. | - `out["keypoints"]`, `out["scores"]` and `out["descriptors"]` also work. These are the same keys as the Hugging Face `SuperPointImageProcessor.post_process_keypoint_detection` result. The keypoints stay as floats here (Hugging Face converts them to int32). - `out.numpy()` gives a dict of numpy arrays. `out.to_dict()` gives a JSON-friendly dict. `len(out)` is N. - The output is equal to the HTTP server response for the same image and parameters. The server rounds the JSON floats (3 decimals for keypoints, 6 for scores) and sends float16 descriptors. ## Lifetime - `model.close()` releases the traces and the device tensors and stops the host threads of list calls. It closes the device if `from_pretrained` opened it. You can call `close()` more than one time. - `with SuperPoint.from_pretrained() as model:` calls `close()` at the end of the block. - A lock serialises the device work. You can share one model object between threads. - One process can have one open model for each chip. Use one process (or one model) for each chip to use more chips. ## Warm-up `from_pretrained` prepares the per-request variants that most users use. The first real call of a prepared variant is as fast as the later calls (within approximately 10 %, see the table below). The warm-up runs synthetic images (drawn shapes and noise, approximately 650 keypoints each, `tt_superpoint.warmup.synthetic_image`) through the full request path. It does not keep any result: every real call runs the full request path, and the outputs are bit-identical to a model without the warm-up. The warm-up prepares these items: - the device programs and the metal trace of the network, the device NMS, the keypoint list and the descriptor sampling; - the host buffers of all 32 readback buckets of the keypoint list (one bucket for each 32 keypoints) and their first-read check; - the exact fallbacks of the keypoint list: more than 1024 NMS candidates (host top-k and the sampling trace) and a candidate-slot overflow (host post-processing of the resident maps); - the device resize trace of each warm-up size; - the device NMS trace of each warm-up radius (1..8), or the host NMS path (radius 0 and 9..32); - the host post-processing (top-k, sort, descriptor normalisation) with arrays of realistic sizes; - the Pillow decoders (plug-in import and decoder set-up) on files of the largest warm-up size. The warm-up also lets Pillow keep freed image buffers for the next decode (`PIL.Image.core.set_blocks_max`, 2 buffers for each decoding thread, at most 8 MB each for 1920x1080). Without them, each decode of a large image gets new memory from the kernel (1-3 ms for a 1600x900 JPEG). If you set `PILLOW_BLOCKS_MAX`, your value is kept; - the host threads of list calls. The thread pool is kept between calls, and `close()` stops it. ### `warmup_variants` | value | prepares | startup (cached kernels) | |---|---|---:| | `None` or `"default"` | sizes 1920x1080, 1600x900, 1280x720 and 640x480; `nms_radius` 4 and 0 (0 prepares the host NMS of radii 0 and 9..32); JPEG and PNG decoders; 4 list workers; the keypoint-list fallbacks | 4.7-5.6 s (the warm-up part: 0.6-0.8 s) | | `"minimal"` or `False` | the default call (radius 4) on a 640x480 image only | 4.2 s | | `"all"` | the default, and every `nms_radius` 0..9 | 6.0 s (16-28 s the first time on a machine, because the kernels of the other radii compile) | | a dict | the default, with the given keys replaced, for example `{"nms_radii": (0, 3, 4)}` or `{"sizes": ((1920, 1080), (1024, 768))}` | | Dict keys: `sizes` (source sizes `(width, height)`), `nms_radii`, `decoders` (`"jpeg"`, `"png"`, `"bmp"`, `"webp"`, `"tiff"`), `num_workers`, `fallbacks` (bool). `model.config["warmup_variants"]` shows the spec in use, and `model.config["warmup_s"]` shows the time of each part. ### `model.warmup(**variant)` Use `model.warmup(...)` to prepare more variants after `from_pretrained`. A variant that is ready is skipped, so you can call it more than one time. It returns the seconds of each step that ran (an empty dict when all were ready). ```python model.warmup(nms_radius=3) # ~120 ms: device NMS trace for radius 3 (+ its fallbacks) model.warmup(size=(1024, 768)) # ~45 ms: device resize trace for 1024x768 source images model.warmup(sizes=[(800, 600), (640, 360)], decoders=["webp"], num_workers=8) ``` ### Costs | item | startup time (cached kernels) | memory | |---|---:|---| | a source size | 35-75 ms (2 synthetic calls) | the source image buffer on the device (W x H bytes, for example 2 MB for 1920x1080) and one small trace | | an `nms_radius` 1..8 other than 4 | approximately 120 ms (with the fallbacks); 1-3 s when its kernels compile the first time on a machine | approximately 42 MB of device DRAM and one trace | | the host NMS (radius 0) | approximately 25 ms | none | | the keypoint-list fallbacks (for each prepared radius) | approximately 25 ms | none | | a decoder | 55 ms (JPEG), 210 ms (PNG): encode a synthetic file at the largest warm-up size, then decode it 3 times | Pillow buffers, see above | | `num_workers` threads | approximately 110 ms (each thread decodes a file of the largest warm-up size) | the threads | ### What the warm-up cannot prepare | case | first-use cost | what to do | |---|---|---| | A source size that is not in `sizes` (the device resize is built on first use) | approximately 10 ms one time; up to approximately 0.2 s on a loaded host; approximately 0.6 s when the resize kernel for its tap counts compiles the first time on a machine | Add the size to `warmup_variants["sizes"]` or call `model.warmup(size=(w, h))`. At most 8 sizes are kept (the least recently used one is released and built again on its next use). | | An `nms_radius` 1..8 that is not prepared | approximately 40 ms one time (1-3 s when its kernels compile the first time) | Add the radius to `nms_radii` or call `model.warmup(nms_radius=r)`. | | Images outside the device resize range (host resize), for example 4032x3024 or 8000x4500 | the first call is approximately 20 % slower (the host allocator grows to the image size); there is no device variant to build | Nothing to prepare: the sizes are not bounded. | | The first `from_pretrained` on a machine (empty kernel cache) | 10-60 s of kernel compile | Run the model one time; the kernel cache (`TT_METAL_CACHE`) keeps the kernels. | The latency of each call depends on the number of keypoints of the image. The model reads a readback bucket that it selects from the keypoint count of the previous image, and reads a second bucket when the guess is wrong. This is a property of the image sequence, not a first-use cost. ### Measured first-call latency `bench_first_call.py`: one fresh process for each variant and repetition, chip with ETH dispatch, 1 command queue and a 12x10 grid, loaded shared host (load average 15-19 on 64 cores), cached kernels. The inputs are 38 real images (19 photographs and their mirror images) at the size of the variant. Every call of pass 1 gets a new image. The reference ("steady") is the median latency of the same image in 3 later passes, because the latency depends on the image. Values are the median of 3 repetitions. Before: commit 588122a. After: this version with the default `warmup_variants`. | variant | `from_pretrained` before / after | call 1 before (ratio to steady) | call 1 after (ratio) | call 2 after (ratio) | call 10 after (ratio) | steady median, all images, before / after | |---|---:|---:|---:|---:|---:|---:| | JPEG file, 1600x900 | 4.10 / 4.95 s | 10.95 ms (1.30) | 9.03 ms (1.07) | 8.75 ms (1.04) | 5.97 ms (0.89) | 6.96 / 7.51 ms | | RGB array, 1920x1080 | 4.18 / 4.72 s | 3.50 ms (1.38) | 2.59 ms (1.04) | 2.51 ms (1.02) | 2.53 ms (1.00) | 2.66 / 2.49 ms | | RGB array, 1280x720 | 4.37 / 5.66 s | 2.31 ms (1.24) | 1.68 ms (0.97) | 1.75 ms (0.91) | 1.78 ms (0.96) | 1.72 / 1.73 ms | | plane, 480x640 | 4.29 / 5.20 s | 2.01 ms (1.64) | 1.13 ms (0.90) | 1.20 ms (0.99) | 1.08 ms (0.92) | 1.21 / 1.24 ms | | PNG file, 1600x900 | 4.39 / 5.62 s | 34.59 ms (1.11) | 32.82 ms (1.02) | 30.70 ms (1.01) | 31.67 ms (1.00) | 31.9 / 31.9 ms | | `max_keypoints=-1`, 1600x900 | 4.36 / 4.94 s | 2.59 ms (1.19) | 2.15 ms (1.01) | 2.13 ms (0.93) | 2.12 ms (0.84) | 2.35 / 2.30 ms | | `return_descriptors=False`, 1600x900 | 4.65 / 4.83 s | 2.35 ms (1.12) | 1.99 ms (1.05) | 1.90 ms (1.02) | 1.94 ms (1.02) | 1.91 / 1.87 ms | | `nms_radius=0` (host NMS), 1600x900 | 4.12 / 4.94 s | 10.05 ms (1.14) | 8.37 ms (1.15) | 7.67 ms (1.05) | 8.41 ms (1.16) | 8.17 / 7.21 ms | | `nms_radius=12` (host NMS), 1600x900 | 4.97 / 5.45 s | 168.0 ms (1.03) | 151.1 ms (0.99) | 151.3 ms (0.99) | 153.6 ms (1.00) | 158.8 / 153.2 ms | | list of 4 JPEG files, 1600x900 | 4.47 / 4.66 s | 27.04 ms (1.25) | 23.51 ms (1.16) | 28.44 ms (1.23) | 24.10 ms (1.20) | 22.8 / 18.7 ms | | `nms_radius=3`, not prepared (default spec) | 4.16 / 5.64 s | 43.81 ms (20.2) | 40.45 ms (16.5) | 2.58 ms (1.18) | 2.56 ms (1.07) | 2.18 / 2.17 ms | | `nms_radius=3`, `warmup_variants={"nms_radii": (0, 3, 4)}` | 4.70 s | | 2.22 ms (1.01) | 2.13 ms (1.01) | 2.17 ms (0.93) | 2.17 ms | | RGB array 1024x768, not prepared (default spec) | 4.16 / 5.20 s | 12.21 ms (6.98) | 11.88 ms (7.22) | 1.72 ms (1.07) | 1.77 ms (1.12) | 1.72 / 1.64 ms | | RGB array 1024x768, `1024x768` added to `sizes` | 4.62 s | | 1.63 ms (0.99) | 1.72 ms (1.01) | 1.72 ms (1.06) | 1.68 ms | | RGB array 8000x4500 (host resize, cannot be prepared) | 4.77 / 5.07 s | 75.40 ms (1.18) | 69.12 ms (1.22) | 57.66 ms (0.99) | 57.67 ms (0.95) | 59.0 / 60.8 ms | - The outputs of all 13 variants (3 repetitions each) are bit-identical before and after the change, and the same image gives the same output in every pass. - The steady latency did not change (the differences are in the noise of the shared host). - The host-NMS rows (radius 0 and 12) and the list row are mostly host work. Their call 2 and call 10 vary as much as call 1 (+-20 % on this host), so the remaining call-1 difference is noise, not a first-use cost. - `models/tests/test_api_warmup_device.py` checks this in one process: the first call of 9 variants (7 from the default spec, 2 added with `model.warmup`) against the steady latency of the same image, and the outputs against a model with `warmup_variants="minimal"`. ## Console output With `verbose=False` (the default), `from_pretrained` sets `TT_LOGGER_LEVEL=Error` (tt-metal C++ log) and `LOGURU_LEVEL=WARNING` (ttnn Python log) before it imports ttnn. Warnings and errors are still shown. The tt-metal info lines (device open, kernel cache statistics) and the warnings that the warm-up causes on purpose (device buffers allocated after a trace capture) are not shown. - A value that you set in the environment is kept. - When ttnn is already imported (for example, when you pass `device=`), the levels cannot change. Set the two variables before you import ttnn. - `verbose=True` keeps the tt-metal defaults (info log). - The Hugging Face `transformers` progress bar of the weight load ("Loading weights") is not changed. ## Examples Match two images (mutual nearest neighbours on the descriptors): ```python import torch from tt_superpoint import SuperPoint with SuperPoint.from_pretrained() as model: a, b = model(["left.jpg", "right.jpg"]) sim = a.descriptors @ b.descriptors.T i = sim.argmax(1); j = sim.argmax(0) mutual = j[i] == torch.arange(len(i)) pairs = torch.stack([a.keypoints[mutual], b.keypoints[i[mutual]]], 1) # (M, 2, 2) xy pairs ``` OpenCV frames: ```python import cv2 frame = cv2.imread("frame.png") # BGR out = model(frame, bgr=True) kps = [cv2.KeyPoint(float(x), float(y), 8) for x, y in out.keypoints.tolist()] ``` ## Performance Warm calls, 1600x900 demo image, chip with ETH dispatch, 1 command queue and a 12x10 grid (the default), loaded shared host (`models/tests/test_api_device.py`). Re-measured on 2026-10-05 in this configuration (chip 3): 1.58 / 1.47, 1.93 / 1.82, 1.15 / 0.95 and 17.6 / 16.1 ms (median / min) for the 4 rows below, the same within the noise of the shared host; first call of the 9 warm-up variants of `test_api_warmup_device.py` 0.91-1.08 of the steady latency. See OPT_REPORT.md "p150 ETH-dispatch compliance 2026-10-05". | call | median | min | |---|---:|---:| | `model(r_plane)`, 1600x900 uint8 R plane (device resize) | 1.59 ms | 1.50 ms | | `model(rgb)`, 1600x900x3 uint8 array (includes the host R-channel copy) | 1.94 ms | 1.85 ms | | `model(plane)`, 480x640 uint8 plane | 1.08 ms | 0.99 ms | | `model("image.jpg")`, includes the JPEG decode (approximately 15 ms) | 16.8 ms | 16.2 ms | The server request path in the same process takes 2.70 / 2.66 / 2.11 / 18.8 ms (median) for the same inputs, because it also builds the JSON lists and the NPZ. The published served numbers for the 1600x900 image are 1.73-1.85 ms (device_forward + postprocess). The published 480x640 request path (`e2e_kpc_u8`, 0.90-0.96 ms) does not include the host sort of the keypoints by score, approximately 0.1 ms of each call.