hamer-p150

HaMeR (Hand Mesh Recovery, ViT-H/16 + MANO regression head) port on one Tenstorrent Blackhole p150a. Weights: changh95/hamer-weights · Paper: arXiv:2312.05251 · Upstream code: geopavlakos/hamer · Port: changh95/tt-hamer

Runs on p150 (mesh P150). Dispatch runs on Ethernet cores with 1 command queue, so all 12×10 Tensix cores compute.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a Python environment with a built tt-metal / ttnn at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied.

hf download changh95/hamer-p150 --exclude "image/*" --local-dir hamer-p150 && cd hamer-p150
pip install -e code/    # host dependencies only; "code/[server]" also installs the HTTP server
pip install -e "code/[server,test]"   # optional: also the HTTP server and the tests (pytest, pyyaml)
from hamer import HamerModel   # pip install -e code/  (in an environment with ttnn / tt-metal)

with HamerModel.from_pretrained(device_id=0) as model:   # weights from the HF cache, opens the chip, warms up
    hand = model("media/sample1.jpg", bbox=[20, 170, 295, 260], is_right=True)
    print(hand.rotmats.shape, hand.betas.shape, hand.cam_t_full)   # (16, 3, 3) (10,) [tx ty tz]

    # several hands in one frame: one box and one is_right per hand
    hands = model("media/sample1.jpg", bbox=[[20, 170, 295, 260], [130, 340, 320, 510]],
                  is_right=[True, False])
  • from_pretrained applies the published serve configuration and downloads the pinned weights changh95/hamer-weights (2.7 GB) to the HF cache.
  • It opens the chip with ETH dispatch, 1 command queue and the 12×10 grid, then runs the warm-up PCC gate. model.info shows dispatch, num_command_queues and grid. Boot takes approximately 10 s with a warm kernel cache.
  • The with block closes the model and the chip at the end. You can also call model.close().
  • from_pretrained() also warms up all call variants. The first call is as fast as the next calls (1.03x for a JPEG file, was 1.34x). Start-up time increases by about 0.25 s.
  • Use warmup_variants=None to skip the host warm-up. Use model.warmup(frame_size=(W, H)) to prepare frames larger than 1920x1080.
  • Run the example from the repository root. The path "media/sample1.jpg" is relative to the repository root.
name meaning
input image full frame: file path, PIL.Image, (H, W, 3) uint8 numpy array or torch tensor
input bbox hand box [x1, y1, x2, y2] in original pixels, or a list of boxes (you get a list)
option is_right True (default); False for a left hand; one value per box
option rescale_factor crop side = factor × longer box side; default 2.0
option return_vertices include the 778 mesh vertices (MANO only); default True
output rotmats (16, 3, 3): index 0 = global orientation, 1..15 = MANO hand pose
output betas, cam (10,) MANO shape; (3,) crop camera [s, tx, ty]
output cam_t_full (3,) camera translation [tx, ty, tz] for the full frame
output vertices, joints, keypoints_2d (778, 3), (21, 3) metres; (21, 2) pixels; None without MANO_RIGHT.pkl
output to_dict() the same JSON as the HTTP /predict response

Throughput helpers:

  • A list of boxes decodes the frame one time and runs one device call for each hand.
  • model.forward_crops(img) takes a batch of crops that you made, (B, 3, 256, 256) or (B, 3, 256, 192). It returns the upstream HAMER.forward keys.
  • model.run_crop(crop) runs one device call on one (1, 3, 256, 192) network input. The "model call" number below measures this call.
  • open_device() and from_pretrained(device=...) let you share one open chip with other code.

To use the MANO mesh, give mano_path= with your own MANO_RIGHT.pkl. PYTHON.md gives the full API reference. examples/quickstart.py runs the snippet on both demo hands.

Serving (HTTP)

tt-model pull  changh95/hamer-p150 --with-weights
tt-model serve changh95/hamer-p150      # or, with tt-cli: tt serve changh95/hamer-p150
printf '{"image":"%s","bbox":[20,170,295,260],"is_right":true}' "$(base64 -w0 media/sample1.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/hamer-p150       # tt-cli
  • The server listens on port 20000 (or the next free port). It is ready when the log shows Application startup complete.
  • POST /predict: image (base64 PNG/JPEG of the full frame) and bbox [x1, y1, x2, y2] (sample boxes: media/bboxes.json). Optional: is_right (true), rescale_factor (2.0), return_vertices (true), return_faces (false).
  • GET /health, GET /info.
{"image_size": [334, 512], "bbox": [20.0, 170.0, 295.0, 260.0], "is_right": true,
 "rotmats": [[[0.957649, 0.103436, 0.268718], "..."], "..."], "betas": [-0.257281, -0.099984, -0.43662, "..."],
 "cam": [4.691406, -0.051758, 0.003906], "cam_t_full": [-0.059121, -0.027873, 7.751116],
 "mano_available": false, "trace_active": true,
 "timing_ms": {"preprocess": 1.19, "device": 3.55, "mano": 0.0, "total": 4.85}}
  • The timing_ms values in the example are the p50 values of the current code/ server (ETH dispatch, 1 command queue, 12×10 grid). The container image is older than code/ (see Provenance). It uses stock worker dispatch, which gives an 11×10 grid on a p150, and it shows timing_ms.device of approximately 13.4 ms until the image is rebuilt.
  • The server boot log shows Device dispatch=eth command_queues=1 compute_grid=12x10. GET /info gives the same three values.
  • The response fields are the same as the Python outputs. cam_t_full uses focal_length = 5000/256*max(W, H). With MANO loaded, the response also has joints, vertices and keypoints_2d, and faces [1538][3] on request.

Demo

Input (media/sample1.jpg, InterHand2.6M) Mesh overlay: input · torch CPU reference · tt-nn on p150a (media/sample1_result.png)

Demo & Performances

Warm, batch 1, one hand crop (256×192) per call, traced fused graph, optimized serve env. scripts/bench.py --iters 100 (3 runs) and scripts/serve_gate.sh (300 HTTP requests per run).

Metric Performance
Model call TtHamer.__call__ (crop prep + H2D + trace + readback + host finalize), fresh crop each call 3.49 ms median (3.483–3.498, 3 runs) · 3.460 ms min
Device trace (ViT-H/16 backbone + MANO head, back-to-back replay) 3.357 ms mean (3.356–3.357, 3 runs)
Host prep · H2D · trace + readback · host finalize 0.031–0.039 · 0.024–0.027 · 3.374–3.377 · 0.047–0.052 ms
Served /predict, timing_ms.device p50 3.55 ms (another run on a busy host: 3.63 ms)
Served /predict, timing_ms.total p50 (host preprocess p50 1.19 ms) 4.85 ms (another run on a busy host: 5.73 ms)
Python model(frame, bbox) call (host crop + resize + model call + post-processing, no MANO), chip 9 4.13–4.14 ms median (2 runs) · 3.980 ms min

The measurement configuration is the p150 configuration: a Blackhole chip with a 12×10 compute grid, dispatch on ETH cores and 1 command queue. No number on this card uses worker dispatch or a second command queue. Accuracy against the torch CPU fp32 reference (scripts/acc.py, 32 real hand crops, served uint8 input): regression-vector PCC mean 0.999993, min 0.999974, 0 crops below the 0.9999 gate; served warm-up gate 0.99970 eager · 1.00000 replay · 0.99985 fresh input (gate 0.99). Details: VERIFICATION_2026-10-03.md.

The Python API row was measured on a different chip (chip 9) on 2026-10-04. On chip 9, model.run_crop is 3.61 ms median and the bare TtHamer call is 3.60 ms. Thus the API adds at most 0.01 ms. Chip 9 is approximately 0.1 ms slower than chip 5, which gave the other rows. hand.to_dict() is equal to the /predict JSON for all 16 test cases. The API does not change the device graph or the numerics.

Re-check 2026-10-05 (p150 ETH-dispatch compliance, commit aaf8601): ETH dispatch, 1 command queue and the 12×10 grid are now the defaults on all paths (Python API, server, bare TtHamer, scripts/bench.py). An independent verifier recorded these values at device open. On chips 3 and 4, the device trace is 3.400–3.425 ms mean, the model call is 3.61–3.65 ms median and the served timing_ms.device p50 is 3.65–3.69 ms. A same-chip A/B against the previous commit shows no change (chip 3: 3.420 ms before, 3.420–3.425 ms after). The difference from the chip-5 rows above is a chip-to-chip difference. Outputs are bit-identical and the accuracy values are unchanged. Details: VERIFICATION_2026-10-03.md, section "p150 ETH-dispatch compliance (2026-10-05)".

RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the sample1 hand. The "incl. H2D/D2H" column compares with our model call (3.49 ms). The "forward only" column compares with our device trace (3.36 ms). Full table: GPU_COMPARISON.md.

RTX 5090 precision GPU incl. H2D/D2H vs current build (3.49 ms) GPU forward only vs ours (3.36 ms)
fp32 strict, eager 13.09 ms Blackhole 3.75× faster 12.93 ms: Blackhole 3.85× faster
tf32, eager 7.42 ms Blackhole 2.12× faster 7.34 ms: Blackhole 2.18× faster
bf16 autocast, eager 7.89 ms Blackhole 2.26× faster 7.60 ms: Blackhole 2.26× faster
fp16 autocast, eager 8.62 ms Blackhole 2.47× faster 8.52 ms: Blackhole 2.54× faster
bf16 weights, eager 5.60 ms Blackhole 1.61× faster 5.86 ms: Blackhole 1.74× faster
bf16 autocast + torch.compile (CUDA graphs) 5.53 ms Blackhole 1.58× faster 5.17 ms: Blackhole 1.54× faster
bf16 weights + torch.compile (CUDA graphs) 3.62 ms Blackhole 1.04× faster 3.39 ms: parity (1.01×)

The device graph is one metal trace of custom fused kernels on all 120 compute cores. These kernels include the ViT-H matmuls with fused add+LN and GELU, and MANO-head GEMVs with weight prefetch. The model call is faster than the best GPU variant because the host path is small: a 295 KB uint8 upload and persistent host buffers. The previous release (13.5 ms device, stock worker dispatch, 11×10 grid) was 3.73× slower than the best GPU variant.

Caveats

  • Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
  • HAMER_DISPATCH=worker (Python: dispatch="worker") is an opt-in that is equivalent only on a Galaxy Blackhole chip. On a p150 it gives an 11×10 grid. The default auto uses this mode only if ETH dispatch does not open, and then gives a RuntimeWarning. No number on this card uses it.
  • One hand per device call: the box is cropped to a 256×192 network input, batch 1. The Python API and forward_crops run one device call for each box. Without bbox, the model uses a centre square and the camera is not correct.
  • bf16 / BFP8 on device: outputs differ slightly from the fp32 reference (PCC above); boot aborts if any warm-up PCC drops below 0.99. Two optimizations change the fp32 summation order: the fused patch matmul and the custom MANO-head attention (kernels/hattn2). OPT_REPORT.md lists each change and its effect.
  • The launch-fused device graph (TT_FUSED=1, default) is what the numbers above measure; TT_FUSED=0 restores the pre-fusion path. The MANO joint metric (scripts/mpjpe.py) was not run, because the measurement host has no MANO_RIGHT.pkl.
  • MANO_RIGHT.pkl is not redistributed: copy your own to ~/.cache/tt-model/hamer-p150/weights/ (or give mano_path= in Python) for mesh, joints and 2-D keypoints; otherwise the server returns MANO parameters only (mano_available: false).
  • The host preprocess compiles a small C library on first use (code/hamer/native/resize.c). Without a C compiler, the server uses the bit-identical PIL path, which is about 0.8 ms slower.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • Measured on tt-metal 8b98410e730 (v0.78.0-dev20260820) with patches/tt-metal-eth-dispatch.patch, ETH dispatch, 1 command queue and the 12×10 compute grid (2026-10-03; re-check 2026-10-05).
  • p150a power was not measured, so no efficiency comparison is made.

Licensing

Provenance

These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build, see OPT_REPORT.md; 2026-10-04 Python API, see PYTHON.md; 2026-10-05 ETH dispatch + 1 command queue + 12×10 as the default on all paths) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:

component built from
tt-metal 8b98410e730bb504fea43a88609756e34821d91d
code/ digest (image) 4156e356e27b902c (sha256, first 16 hex digits; the current code/ differs)
built 2026-09-13T15:34:00+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including changh95/hamer-p150

Paper for changh95/hamer-p150