superpoint-p150 / code /PYTHON.md
changh95's picture
p150 ETH-dispatch compliance (2026-10-05): default ETH dispatch, 1 CQ, 12x10 in Python API and server; numbers re-measured
026da6c verified
|
Raw History Blame Contribute Delete
21.3 kB

tt_superpoint: Python API

Use SuperPoint (keypoint detection and 256-d descriptors) on one Tenstorrent Blackhole chip from Python code, notebooks and pipelines. The HTTP server (models/server/app.py, see SERVING.md) runs the same request path.

Install

Install the package into a Python environment that has ttnn (tt-metal). Run the commands from the repo root:

pip install -e code/                    # the Python API only (or: pip install code/)
pip install -e "code/[server,test]"     # also the HTTP server (fastapi, uvicorn, pydantic) and pytest

The package declares only its own dependencies (numpy, torch, pillow, transformers, huggingface_hub, safetensors). ttnn is not a PyPI dependency. You do not need PYTHONPATH or sys.path changes.

to run extras needs
the Python API none ttnn and a chip
the HTTP server (models/server/app.py, see SERVING.md) server ttnn and a chip
the host tests (no chip) server,test ttnn importable (the tests compare with the server code and import the port modules), no device
the device tests server,test ttnn and a chip

Run only the host tests (no device is opened):

cd code && TT_VISIBLE_DEVICES=none python -m pytest -q models/tests/test_api_host.py models/tests/test_api_warmup_host.py

The device tests of the API are models/tests/test_api_device.py and models/tests/test_api_warmup_device.py (cd code && python -m pytest -s -q <file>).

The weights (magic-leap-community/superpoint, 5 MB) come from the Hugging Face cache. The first call downloads them if they are not in the cache.

Quickstart

from tt_superpoint import SuperPoint

with SuperPoint.from_pretrained(device_id=0) as model:
    out = model("code/sample_data/house_in_field_1080p.jpg")
    print(len(out), out.keypoints[:3], out.scores[:3], out.descriptors.shape)

The relative path code/sample_data/... in the snippets of this file assumes that the current directory is the repo root. From another directory, use an absolute path.

python code/examples/quickstart.py runs this snippet on the demo image. The script finds the demo image from its own location, so you can run it from any directory. It writes keypoints.png, keypoints.json and descriptors.npy to quickstart_out/ in the current directory (--out-dir changes this).

SuperPoint.from_pretrained(...)

SuperPoint.from_pretrained(
    pretrained_model_name_or_path="magic-leap-community/superpoint",
    *, revision=None, device=None, device_id=0, dispatch="auto",
    warmup_variants=None,              # None = "default"; see "Warm-up" below
    precompile_sizes=None, precompile_nms_radii=(),
    trace_region_size=32 << 20, local_files_only=False, verbose=False,
) -> SuperPoint

This call loads the weights, opens the chip, builds the model, compiles the kernels, captures the metal traces and warms up the per-request variants (see Warm-up). Startup takes approximately 5 s with cached kernels and 10-60 s without them.

argument meaning
pretrained_model_name_or_path Hugging Face repo id, or a local directory with config.json and model.safetensors.
revision Weights revision. The default for the default repo is 734450e9, the revision of the published numbers.
device An open ttnn device. The device must have l1_small_size >= 32768 and trace_region_size >= 32 MiB. close() does not close this device.
device_id The chip to open when device is None.
dispatch "auto" (default): ETH dispatch, 1 command queue and a 12x10 compute grid when the tt-metal tree has the ETH-dispatch patch. This is the p150 target configuration and the configuration of all published numbers. Without the patch, the default (Tensix) dispatch is used and a RuntimeWarning is shown; on a p150 the compute grid is then 11x10. The outputs are the same. "eth" forces ETH dispatch. "worker" forces the Tensix dispatch: an explicit opt-in for comparisons on a Galaxy chip (12x10 there, 11x10 on a p150). $SP_DISPATCH overrides "auto". The model uses only command queue 0.
warmup_variants The per-request variants to prepare at startup, so that the first real call of each one is as fast as the later calls. See Warm-up.
precompile_sizes Older name. When given, it replaces warmup_variants["sizes"].
precompile_nms_radii Older name. The values are added to warmup_variants["nms_radii"].
trace_region_size Trace region in bytes when this call opens the device.
local_files_only Load the weights from the Hugging Face cache only (no network).
verbose False (default): tt-metal and ttnn show only warnings and errors. True: they also show their info log. See Console output.

model(images, **params)

out  = model(image, max_keypoints=1024, keypoint_threshold=0.005, nms_radius=4,
             return_descriptors=True, bgr=False)          # -> SuperPointOutput
outs = model([image1, image2, ...], num_workers=4)         # -> list[SuperPointOutput]

Images:

input notes
str / pathlib.Path An image file (PNG, JPEG, or any format that Pillow opens).
bytes Encoded file contents.
PIL.Image.Image Any mode.
numpy.ndarray (H, W) gray, or (H, W, C) with C = 1, 3 or 4 (RGB or RGBA). Use bgr=True for OpenCV (cv2.imread) arrays.
torch.Tensor (H, W), (C, H, W) (torchvision) or (H, W, C).
a list or tuple One output for each image, in input order.
4-D array or tensor A batch: (B, H, W, C) numpy or (B, C, H, W) torch. One output for each image.
  • dtypes: uint8 (0..255), other integers in 0..255, or floats in [0, 1]. The model rounds float values to 8 bits, because the device reads 8-bit images.
  • Any image size is accepted. The model resizes every image to the 480x640 network frame (bilinear, the aspect ratio is not kept). The resize is bit-exact with Pillow. It runs on the device when the size fits the device resize kernel, for example 3840x2160 and the smaller common sizes (one resize trace for each size, built at startup for the warm-up sizes and on first use for other sizes, 8 sizes kept). Other sizes are resized on the host, for example 4032x3024 (12 MP phone photos).
  • The model reads channel 0 (R) of the RGB image, like the Hugging Face SuperPointForKeypointDetection. For a grayscale image, channel 0 is the gray plane.

Parameters (the same defaults and limits as the HTTP server):

parameter default range effect
max_keypoints 1024 -1..307200 Keep the top-k keypoints by score. -1 keeps all keypoints above the threshold.
keypoint_threshold 0.005 [0, 1] Minimum score after NMS.
nms_radius 4 0..32 NMS radius in network-frame pixels. Radii 1..8 run on the device (4 is in the main trace; other radii have a separate trace). Radius 0 (no NMS) and radii above 8 run on the host, which is slower.
return_descriptors True bool When False, descriptors is None.
bgr False bool Arrays are in BGR order (OpenCV).
num_workers 4 int >= 0 Lists only: host threads that decode the next images while the device runs the current image. 0 = no threads.

model.iter(images, **params) gives the same outputs as model(list) one at a time. It accepts any iterable, for example a generator over video frames.

SuperPointOutput

field type content
keypoints torch.float32 (N, 2) [x, y] in pixels of the original image (network-frame position multiplied by scale).
scores torch.float32 (N,) Keypoint scores in (0, 1], sorted in descending order. keypoints and descriptors use the same order.
descriptors torch.float32 (N, 256) or None L2-normalised descriptors, bilinearly sampled at the keypoints.
image_size (int, int) (height, width) of the original image.
scale (float, float) (sx, sy) = original size / (640, 480).
device_nms bool True when NMS ran on the device.
timing_ms dict preprocess, device_forward, postprocess of this call.
  • out["keypoints"], out["scores"] and out["descriptors"] also work. These are the same keys as the Hugging Face SuperPointImageProcessor.post_process_keypoint_detection result. The keypoints stay as floats here (Hugging Face converts them to int32).
  • out.numpy() gives a dict of numpy arrays. out.to_dict() gives a JSON-friendly dict. len(out) is N.
  • The output is equal to the HTTP server response for the same image and parameters. The server rounds the JSON floats (3 decimals for keypoints, 6 for scores) and sends float16 descriptors.

Lifetime

  • model.close() releases the traces and the device tensors and stops the host threads of list calls. It closes the device if from_pretrained opened it. You can call close() more than one time.
  • with SuperPoint.from_pretrained() as model: calls close() at the end of the block.
  • A lock serialises the device work. You can share one model object between threads.
  • One process can have one open model for each chip. Use one process (or one model) for each chip to use more chips.

Warm-up

from_pretrained prepares the per-request variants that most users use. The first real call of a prepared variant is as fast as the later calls (within approximately 10 %, see the table below). The warm-up runs synthetic images (drawn shapes and noise, approximately 650 keypoints each, tt_superpoint.warmup.synthetic_image) through the full request path. It does not keep any result: every real call runs the full request path, and the outputs are bit-identical to a model without the warm-up.

The warm-up prepares these items:

  • the device programs and the metal trace of the network, the device NMS, the keypoint list and the descriptor sampling;
  • the host buffers of all 32 readback buckets of the keypoint list (one bucket for each 32 keypoints) and their first-read check;
  • the exact fallbacks of the keypoint list: more than 1024 NMS candidates (host top-k and the sampling trace) and a candidate-slot overflow (host post-processing of the resident maps);
  • the device resize trace of each warm-up size;
  • the device NMS trace of each warm-up radius (1..8), or the host NMS path (radius 0 and 9..32);
  • the host post-processing (top-k, sort, descriptor normalisation) with arrays of realistic sizes;
  • the Pillow decoders (plug-in import and decoder set-up) on files of the largest warm-up size. The warm-up also lets Pillow keep freed image buffers for the next decode (PIL.Image.core.set_blocks_max, 2 buffers for each decoding thread, at most 8 MB each for 1920x1080). Without them, each decode of a large image gets new memory from the kernel (1-3 ms for a 1600x900 JPEG). If you set PILLOW_BLOCKS_MAX, your value is kept;
  • the host threads of list calls. The thread pool is kept between calls, and close() stops it.

warmup_variants

value prepares startup (cached kernels)
None or "default" sizes 1920x1080, 1600x900, 1280x720 and 640x480; nms_radius 4 and 0 (0 prepares the host NMS of radii 0 and 9..32); JPEG and PNG decoders; 4 list workers; the keypoint-list fallbacks 4.7-5.6 s (the warm-up part: 0.6-0.8 s)
"minimal" or False the default call (radius 4) on a 640x480 image only 4.2 s
"all" the default, and every nms_radius 0..9 6.0 s (16-28 s the first time on a machine, because the kernels of the other radii compile)
a dict the default, with the given keys replaced, for example {"nms_radii": (0, 3, 4)} or {"sizes": ((1920, 1080), (1024, 768))}

Dict keys: sizes (source sizes (width, height)), nms_radii, decoders ("jpeg", "png", "bmp", "webp", "tiff"), num_workers, fallbacks (bool). model.config["warmup_variants"] shows the spec in use, and model.config["warmup_s"] shows the time of each part.

model.warmup(**variant)

Use model.warmup(...) to prepare more variants after from_pretrained. A variant that is ready is skipped, so you can call it more than one time. It returns the seconds of each step that ran (an empty dict when all were ready).

model.warmup(nms_radius=3)                 # ~120 ms: device NMS trace for radius 3 (+ its fallbacks)
model.warmup(size=(1024, 768))             # ~45 ms: device resize trace for 1024x768 source images
model.warmup(sizes=[(800, 600), (640, 360)], decoders=["webp"], num_workers=8)

Costs

item startup time (cached kernels) memory
a source size 35-75 ms (2 synthetic calls) the source image buffer on the device (W x H bytes, for example 2 MB for 1920x1080) and one small trace
an nms_radius 1..8 other than 4 approximately 120 ms (with the fallbacks); 1-3 s when its kernels compile the first time on a machine approximately 42 MB of device DRAM and one trace
the host NMS (radius 0) approximately 25 ms none
the keypoint-list fallbacks (for each prepared radius) approximately 25 ms none
a decoder 55 ms (JPEG), 210 ms (PNG): encode a synthetic file at the largest warm-up size, then decode it 3 times Pillow buffers, see above
num_workers threads approximately 110 ms (each thread decodes a file of the largest warm-up size) the threads

What the warm-up cannot prepare

case first-use cost what to do
A source size that is not in sizes (the device resize is built on first use) approximately 10 ms one time; up to approximately 0.2 s on a loaded host; approximately 0.6 s when the resize kernel for its tap counts compiles the first time on a machine Add the size to warmup_variants["sizes"] or call model.warmup(size=(w, h)). At most 8 sizes are kept (the least recently used one is released and built again on its next use).
An nms_radius 1..8 that is not prepared approximately 40 ms one time (1-3 s when its kernels compile the first time) Add the radius to nms_radii or call model.warmup(nms_radius=r).
Images outside the device resize range (host resize), for example 4032x3024 or 8000x4500 the first call is approximately 20 % slower (the host allocator grows to the image size); there is no device variant to build Nothing to prepare: the sizes are not bounded.
The first from_pretrained on a machine (empty kernel cache) 10-60 s of kernel compile Run the model one time; the kernel cache (TT_METAL_CACHE) keeps the kernels.

The latency of each call depends on the number of keypoints of the image. The model reads a readback bucket that it selects from the keypoint count of the previous image, and reads a second bucket when the guess is wrong. This is a property of the image sequence, not a first-use cost.

Measured first-call latency

bench_first_call.py: one fresh process for each variant and repetition, chip with ETH dispatch, 1 command queue and a 12x10 grid, loaded shared host (load average 15-19 on 64 cores), cached kernels. The inputs are 38 real images (19 photographs and their mirror images) at the size of the variant. Every call of pass 1 gets a new image. The reference ("steady") is the median latency of the same image in 3 later passes, because the latency depends on the image. Values are the median of 3 repetitions. Before: commit 588122a. After: this version with the default warmup_variants.

variant from_pretrained before / after call 1 before (ratio to steady) call 1 after (ratio) call 2 after (ratio) call 10 after (ratio) steady median, all images, before / after
JPEG file, 1600x900 4.10 / 4.95 s 10.95 ms (1.30) 9.03 ms (1.07) 8.75 ms (1.04) 5.97 ms (0.89) 6.96 / 7.51 ms
RGB array, 1920x1080 4.18 / 4.72 s 3.50 ms (1.38) 2.59 ms (1.04) 2.51 ms (1.02) 2.53 ms (1.00) 2.66 / 2.49 ms
RGB array, 1280x720 4.37 / 5.66 s 2.31 ms (1.24) 1.68 ms (0.97) 1.75 ms (0.91) 1.78 ms (0.96) 1.72 / 1.73 ms
plane, 480x640 4.29 / 5.20 s 2.01 ms (1.64) 1.13 ms (0.90) 1.20 ms (0.99) 1.08 ms (0.92) 1.21 / 1.24 ms
PNG file, 1600x900 4.39 / 5.62 s 34.59 ms (1.11) 32.82 ms (1.02) 30.70 ms (1.01) 31.67 ms (1.00) 31.9 / 31.9 ms
max_keypoints=-1, 1600x900 4.36 / 4.94 s 2.59 ms (1.19) 2.15 ms (1.01) 2.13 ms (0.93) 2.12 ms (0.84) 2.35 / 2.30 ms
return_descriptors=False, 1600x900 4.65 / 4.83 s 2.35 ms (1.12) 1.99 ms (1.05) 1.90 ms (1.02) 1.94 ms (1.02) 1.91 / 1.87 ms
nms_radius=0 (host NMS), 1600x900 4.12 / 4.94 s 10.05 ms (1.14) 8.37 ms (1.15) 7.67 ms (1.05) 8.41 ms (1.16) 8.17 / 7.21 ms
nms_radius=12 (host NMS), 1600x900 4.97 / 5.45 s 168.0 ms (1.03) 151.1 ms (0.99) 151.3 ms (0.99) 153.6 ms (1.00) 158.8 / 153.2 ms
list of 4 JPEG files, 1600x900 4.47 / 4.66 s 27.04 ms (1.25) 23.51 ms (1.16) 28.44 ms (1.23) 24.10 ms (1.20) 22.8 / 18.7 ms
nms_radius=3, not prepared (default spec) 4.16 / 5.64 s 43.81 ms (20.2) 40.45 ms (16.5) 2.58 ms (1.18) 2.56 ms (1.07) 2.18 / 2.17 ms
nms_radius=3, warmup_variants={"nms_radii": (0, 3, 4)} 4.70 s 2.22 ms (1.01) 2.13 ms (1.01) 2.17 ms (0.93) 2.17 ms
RGB array 1024x768, not prepared (default spec) 4.16 / 5.20 s 12.21 ms (6.98) 11.88 ms (7.22) 1.72 ms (1.07) 1.77 ms (1.12) 1.72 / 1.64 ms
RGB array 1024x768, 1024x768 added to sizes 4.62 s 1.63 ms (0.99) 1.72 ms (1.01) 1.72 ms (1.06) 1.68 ms
RGB array 8000x4500 (host resize, cannot be prepared) 4.77 / 5.07 s 75.40 ms (1.18) 69.12 ms (1.22) 57.66 ms (0.99) 57.67 ms (0.95) 59.0 / 60.8 ms
  • The outputs of all 13 variants (3 repetitions each) are bit-identical before and after the change, and the same image gives the same output in every pass.
  • The steady latency did not change (the differences are in the noise of the shared host).
  • The host-NMS rows (radius 0 and 12) and the list row are mostly host work. Their call 2 and call 10 vary as much as call 1 (+-20 % on this host), so the remaining call-1 difference is noise, not a first-use cost.
  • models/tests/test_api_warmup_device.py checks this in one process: the first call of 9 variants (7 from the default spec, 2 added with model.warmup) against the steady latency of the same image, and the outputs against a model with warmup_variants="minimal".

Console output

With verbose=False (the default), from_pretrained sets TT_LOGGER_LEVEL=Error (tt-metal C++ log) and LOGURU_LEVEL=WARNING (ttnn Python log) before it imports ttnn. Warnings and errors are still shown. The tt-metal info lines (device open, kernel cache statistics) and the warnings that the warm-up causes on purpose (device buffers allocated after a trace capture) are not shown.

  • A value that you set in the environment is kept.
  • When ttnn is already imported (for example, when you pass device=), the levels cannot change. Set the two variables before you import ttnn.
  • verbose=True keeps the tt-metal defaults (info log).
  • The Hugging Face transformers progress bar of the weight load ("Loading weights") is not changed.

Examples

Match two images (mutual nearest neighbours on the descriptors):

import torch
from tt_superpoint import SuperPoint

with SuperPoint.from_pretrained() as model:
    a, b = model(["left.jpg", "right.jpg"])
sim = a.descriptors @ b.descriptors.T
i = sim.argmax(1); j = sim.argmax(0)
mutual = j[i] == torch.arange(len(i))
pairs = torch.stack([a.keypoints[mutual], b.keypoints[i[mutual]]], 1)   # (M, 2, 2) xy pairs

OpenCV frames:

import cv2
frame = cv2.imread("frame.png")          # BGR
out = model(frame, bgr=True)
kps = [cv2.KeyPoint(float(x), float(y), 8) for x, y in out.keypoints.tolist()]

Performance

Warm calls, 1600x900 demo image, chip with ETH dispatch, 1 command queue and a 12x10 grid (the default), loaded shared host (models/tests/test_api_device.py). Re-measured on 2026-10-05 in this configuration (chip 3): 1.58 / 1.47, 1.93 / 1.82, 1.15 / 0.95 and 17.6 / 16.1 ms (median / min) for the 4 rows below, the same within the noise of the shared host; first call of the 9 warm-up variants of test_api_warmup_device.py 0.91-1.08 of the steady latency. See OPT_REPORT.md "p150 ETH-dispatch compliance 2026-10-05".

call median min
model(r_plane), 1600x900 uint8 R plane (device resize) 1.59 ms 1.50 ms
model(rgb), 1600x900x3 uint8 array (includes the host R-channel copy) 1.94 ms 1.85 ms
model(plane), 480x640 uint8 plane 1.08 ms 0.99 ms
model("image.jpg"), includes the JPEG decode (approximately 15 ms) 16.8 ms 16.2 ms

The server request path in the same process takes 2.70 / 2.66 / 2.11 / 18.8 ms (median) for the same inputs, because it also builds the JSON lists and the NPZ. The published served numbers for the 1600x900 image are 1.73-1.85 ms (device_forward + postprocess). The published 480x640 request path (e2e_kpc_u8, 0.90-0.96 ms) does not include the host sort of the keypoints by score, approximately 0.1 ms of each call.