hamer-p150
HaMeR (Hand Mesh Recovery, ViT-H/16 + MANO regression head) port on one Tenstorrent Blackhole p150a. Weights: changh95/hamer-weights · Paper: arXiv:2312.05251 · Upstream code: geopavlakos/hamer · Port: changh95/tt-hamer
Runs on p150 (mesh P150). Dispatch runs on Ethernet cores with 1 command queue, so all 12×10 Tensix cores compute.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a Python environment with a built tt-metal / ttnn at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied.
hf download changh95/hamer-p150 --exclude "image/*" --local-dir hamer-p150 && cd hamer-p150
pip install -e code/ # host dependencies only; "code/[server]" also installs the HTTP server
pip install -e "code/[server,test]" # optional: also the HTTP server and the tests (pytest, pyyaml)
from hamer import HamerModel # pip install -e code/ (in an environment with ttnn / tt-metal)
with HamerModel.from_pretrained(device_id=0) as model: # weights from the HF cache, opens the chip, warms up
hand = model("media/sample1.jpg", bbox=[20, 170, 295, 260], is_right=True)
print(hand.rotmats.shape, hand.betas.shape, hand.cam_t_full) # (16, 3, 3) (10,) [tx ty tz]
# several hands in one frame: one box and one is_right per hand
hands = model("media/sample1.jpg", bbox=[[20, 170, 295, 260], [130, 340, 320, 510]],
is_right=[True, False])
from_pretrainedapplies the published serve configuration and downloads the pinned weightschangh95/hamer-weights(2.7 GB) to the HF cache.- It opens the chip with ETH dispatch, 1 command queue and the 12×10 grid, then runs the warm-up PCC gate.
model.infoshowsdispatch,num_command_queuesandgrid. Boot takes approximately 10 s with a warm kernel cache. - The
withblock closes the model and the chip at the end. You can also callmodel.close(). from_pretrained()also warms up all call variants. The first call is as fast as the next calls (1.03x for a JPEG file, was 1.34x). Start-up time increases by about 0.25 s.- Use
warmup_variants=Noneto skip the host warm-up. Usemodel.warmup(frame_size=(W, H))to prepare frames larger than 1920x1080. - Run the example from the repository root. The path
"media/sample1.jpg"is relative to the repository root.
| name | meaning | |
|---|---|---|
| input | image |
full frame: file path, PIL.Image, (H, W, 3) uint8 numpy array or torch tensor |
| input | bbox |
hand box [x1, y1, x2, y2] in original pixels, or a list of boxes (you get a list) |
| option | is_right |
True (default); False for a left hand; one value per box |
| option | rescale_factor |
crop side = factor × longer box side; default 2.0 |
| option | return_vertices |
include the 778 mesh vertices (MANO only); default True |
| output | rotmats |
(16, 3, 3): index 0 = global orientation, 1..15 = MANO hand pose |
| output | betas, cam |
(10,) MANO shape; (3,) crop camera [s, tx, ty] |
| output | cam_t_full |
(3,) camera translation [tx, ty, tz] for the full frame |
| output | vertices, joints, keypoints_2d |
(778, 3), (21, 3) metres; (21, 2) pixels; None without MANO_RIGHT.pkl |
| output | to_dict() |
the same JSON as the HTTP /predict response |
Throughput helpers:
- A list of boxes decodes the frame one time and runs one device call for each hand.
model.forward_crops(img)takes a batch of crops that you made,(B, 3, 256, 256)or(B, 3, 256, 192). It returns the upstreamHAMER.forwardkeys.model.run_crop(crop)runs one device call on one(1, 3, 256, 192)network input. The "model call" number below measures this call.open_device()andfrom_pretrained(device=...)let you share one open chip with other code.
To use the MANO mesh, give mano_path= with your own MANO_RIGHT.pkl. PYTHON.md gives the full API reference. examples/quickstart.py runs the snippet on both demo hands.
Serving (HTTP)
tt-model pull changh95/hamer-p150 --with-weights
tt-model serve changh95/hamer-p150 # or, with tt-cli: tt serve changh95/hamer-p150
printf '{"image":"%s","bbox":[20,170,295,260],"is_right":true}' "$(base64 -w0 media/sample1.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/hamer-p150 # tt-cli
- The server listens on port 20000 (or the next free port). It is ready when the log shows
Application startup complete. POST /predict:image(base64 PNG/JPEG of the full frame) andbbox[x1, y1, x2, y2](sample boxes:media/bboxes.json). Optional:is_right(true),rescale_factor(2.0),return_vertices(true),return_faces(false).GET /health,GET /info.
{"image_size": [334, 512], "bbox": [20.0, 170.0, 295.0, 260.0], "is_right": true,
"rotmats": [[[0.957649, 0.103436, 0.268718], "..."], "..."], "betas": [-0.257281, -0.099984, -0.43662, "..."],
"cam": [4.691406, -0.051758, 0.003906], "cam_t_full": [-0.059121, -0.027873, 7.751116],
"mano_available": false, "trace_active": true,
"timing_ms": {"preprocess": 1.19, "device": 3.55, "mano": 0.0, "total": 4.85}}
- The
timing_msvalues in the example are the p50 values of the currentcode/server (ETH dispatch, 1 command queue, 12×10 grid). The container image is older thancode/(see Provenance). It uses stock worker dispatch, which gives an 11×10 grid on a p150, and it showstiming_ms.deviceof approximately 13.4 ms until the image is rebuilt. - The server boot log shows
Device dispatch=eth command_queues=1 compute_grid=12x10.GET /infogives the same three values. - The response fields are the same as the Python outputs.
cam_t_fullusesfocal_length = 5000/256*max(W, H). With MANO loaded, the response also hasjoints,verticesandkeypoints_2d, andfaces[1538][3]on request.
Demo
Input (media/sample1.jpg, InterHand2.6M) |
Mesh overlay: input · torch CPU reference · tt-nn on p150a (media/sample1_result.png) |
|---|---|
![]() |
![]() |
Demo & Performances
Warm, batch 1, one hand crop (256×192) per call, traced fused graph, optimized serve env. scripts/bench.py --iters 100 (3 runs) and scripts/serve_gate.sh (300 HTTP requests per run).
| Metric | Performance |
|---|---|
Model call TtHamer.__call__ (crop prep + H2D + trace + readback + host finalize), fresh crop each call |
3.49 ms median (3.483–3.498, 3 runs) · 3.460 ms min |
| Device trace (ViT-H/16 backbone + MANO head, back-to-back replay) | 3.357 ms mean (3.356–3.357, 3 runs) |
| Host prep · H2D · trace + readback · host finalize | 0.031–0.039 · 0.024–0.027 · 3.374–3.377 · 0.047–0.052 ms |
Served /predict, timing_ms.device p50 |
3.55 ms (another run on a busy host: 3.63 ms) |
Served /predict, timing_ms.total p50 (host preprocess p50 1.19 ms) |
4.85 ms (another run on a busy host: 5.73 ms) |
Python model(frame, bbox) call (host crop + resize + model call + post-processing, no MANO), chip 9 |
4.13–4.14 ms median (2 runs) · 3.980 ms min |
The measurement configuration is the p150 configuration: a Blackhole chip with a 12×10 compute grid, dispatch on ETH cores and 1 command queue. No number on this card uses worker dispatch or a second command queue. Accuracy against the torch CPU fp32 reference (scripts/acc.py, 32 real hand crops, served uint8 input): regression-vector PCC mean 0.999993, min 0.999974, 0 crops below the 0.9999 gate; served warm-up gate 0.99970 eager · 1.00000 replay · 0.99985 fresh input (gate 0.99). Details: VERIFICATION_2026-10-03.md.
The Python API row was measured on a different chip (chip 9) on 2026-10-04. On chip 9, model.run_crop is 3.61 ms median and the bare TtHamer call is 3.60 ms. Thus the API adds at most 0.01 ms. Chip 9 is approximately 0.1 ms slower than chip 5, which gave the other rows. hand.to_dict() is equal to the /predict JSON for all 16 test cases. The API does not change the device graph or the numerics.
Re-check 2026-10-05 (p150 ETH-dispatch compliance, commit aaf8601): ETH dispatch, 1 command queue and the 12×10 grid are now the defaults on all paths (Python API, server, bare TtHamer, scripts/bench.py). An independent verifier recorded these values at device open. On chips 3 and 4, the device trace is 3.400–3.425 ms mean, the model call is 3.61–3.65 ms median and the served timing_ms.device p50 is 3.65–3.69 ms. A same-chip A/B against the previous commit shows no change (chip 3: 3.420 ms before, 3.420–3.425 ms after). The difference from the chip-5 rows above is a chip-to-chip difference. Outputs are bit-identical and the accuracy values are unchanged. Details: VERIFICATION_2026-10-03.md, section "p150 ETH-dispatch compliance (2026-10-05)".
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the sample1 hand. The "incl. H2D/D2H" column compares with our model call (3.49 ms). The "forward only" column compares with our device trace (3.36 ms). Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU incl. H2D/D2H | vs current build (3.49 ms) | GPU forward only vs ours (3.36 ms) |
|---|---|---|---|
| fp32 strict, eager | 13.09 ms | Blackhole 3.75× faster | 12.93 ms: Blackhole 3.85× faster |
| tf32, eager | 7.42 ms | Blackhole 2.12× faster | 7.34 ms: Blackhole 2.18× faster |
| bf16 autocast, eager | 7.89 ms | Blackhole 2.26× faster | 7.60 ms: Blackhole 2.26× faster |
| fp16 autocast, eager | 8.62 ms | Blackhole 2.47× faster | 8.52 ms: Blackhole 2.54× faster |
| bf16 weights, eager | 5.60 ms | Blackhole 1.61× faster | 5.86 ms: Blackhole 1.74× faster |
bf16 autocast + torch.compile (CUDA graphs) |
5.53 ms | Blackhole 1.58× faster | 5.17 ms: Blackhole 1.54× faster |
bf16 weights + torch.compile (CUDA graphs) |
3.62 ms | Blackhole 1.04× faster | 3.39 ms: parity (1.01×) |
The device graph is one metal trace of custom fused kernels on all 120 compute cores. These kernels include the ViT-H matmuls with fused add+LN and GELU, and MANO-head GEMVs with weight prefetch. The model call is faster than the best GPU variant because the host path is small: a 295 KB uint8 upload and persistent host buffers. The previous release (13.5 ms device, stock worker dispatch, 11×10 grid) was 3.73× slower than the best GPU variant.
Caveats
- Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (
patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication. HAMER_DISPATCH=worker(Python:dispatch="worker") is an opt-in that is equivalent only on a Galaxy Blackhole chip. On a p150 it gives an 11×10 grid. The defaultautouses this mode only if ETH dispatch does not open, and then gives aRuntimeWarning. No number on this card uses it.- One hand per device call: the box is cropped to a 256×192 network input, batch 1. The Python API and
forward_cropsrun one device call for each box. Withoutbbox, the model uses a centre square and the camera is not correct. - bf16 / BFP8 on device: outputs differ slightly from the fp32 reference (PCC above); boot aborts if any warm-up PCC drops below 0.99. Two optimizations change the fp32 summation order: the fused patch matmul and the custom MANO-head attention (
kernels/hattn2).OPT_REPORT.mdlists each change and its effect. - The launch-fused device graph (
TT_FUSED=1, default) is what the numbers above measure;TT_FUSED=0restores the pre-fusion path. The MANO joint metric (scripts/mpjpe.py) was not run, because the measurement host has noMANO_RIGHT.pkl. MANO_RIGHT.pklis not redistributed: copy your own to~/.cache/tt-model/hamer-p150/weights/(or givemano_path=in Python) for mesh, joints and 2-D keypoints; otherwise the server returns MANO parameters only (mano_available: false).- The host preprocess compiles a small C library on first use (
code/hamer/native/resize.c). Without a C compiler, the server uses the bit-identical PIL path, which is about 0.8 ms slower. - Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - Measured on tt-metal
8b98410e730(v0.78.0-dev20260820) withpatches/tt-metal-eth-dispatch.patch, ETH dispatch, 1 command queue and the 12×10 compute grid (2026-10-03; re-check 2026-10-05). - p150a power was not measured, so no efficiency comparison is made.
Licensing
- Weights: changh95/hamer-weights, the HaMeR authors' released checkpoint mirrored unchanged (MIT code); the MANO hand model it needs is non-commercial research only — do not use this model commercially.
- Port and serving code (
code/): from changh95/tt-hamer, published under the upstream terms (mit-code-mano-non-commercial).
Provenance
These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build, see OPT_REPORT.md; 2026-10-04 Python API, see PYTHON.md; 2026-10-05 ETH dispatch + 1 command queue + 12×10 as the default on all paths) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
4156e356e27b902c (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-13T15:34:00+00:00 by tt-model 0.1.0 |

