moge-2-p150
Microsoft MoGe-2 (ViT-L/normal checkpoint) port on one Tenstorrent Blackhole p150a. Weights: Ruicheng/moge-2-vitl-normal · Paper: arXiv:2507.02546 · Upstream code: microsoft/MoGe · Port: changh95/tt-MoGe
Runs on p150 (mesh P150).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/moge-2-p150 --exclude "image/*" --local-dir moge-2-p150 && cd moge-2-p150
pip install -e code/ # adds numpy<2, pillow, scipy, huggingface_hub, opencv-python-headless
pip install -e "code/[server,test]" # optional: also the HTTP server (fastapi, uvicorn, pydantic, isal) and pytest
Run the commands from the repo root (the directory that contains code/ and media/). The relative paths in this section (code/..., media/...) start at the repo root.
from tt_moge import MoGeModel
with MoGeModel.from_pretrained() as model: # opens device 0, gets the weights from the HF cache
out = model("photo.jpg") # a path, PIL image, numpy array or torch tensor
print(out) # MoGeOutput(1920x1080, valid 90.5%, median depth 13.322 m, fov_x 86.1 deg, metric_scale 10.3329)
depth = out.depth # (H, W) float32, metres; inf where out.mask is False
points = out.points # (H, W, 3) float32, metres, camera space (x right, y down, z forward)
normal = out.normal # (H, W, 3) float32, unit normals
K = out.intrinsics_pixels # (3, 3) camera matrix in pixels
from_pretraineddownloads theRuicheng/moge-2-vitl-normalweights to your HF cache, compiles the kernels and captures the device trace. The first load takes about 40 s. A load with a warm kernel cache takes about 8 s.from_pretrained()also warms up the common call types. Your first call is then as fast as later calls (about 25 ms for a 1920x1080 image). To warm up other input types or sizes, usewarmup_variants=ormodel.warmup(...)(see PYTHON.md).- Load the model one time and call it many times. A warm call on a 1920×1080 image takes about 25 ms.
- The
withblock releases the trace and closes the device at the end. Withoutwith, callmodel.close(). python code/examples/quickstart.pyruns this snippet onmedia/source.png. It writesdepth.png,normal.png,result.npzandresult.jsonto./moge_output/.
Input (image) |
A file path, a PIL.Image, a numpy array (HxWx3 uint8, or float in [0, 1]; also 3xHxW and grayscale), or a torch tensor (3xHxW / 1x3xHxW float in [0, 1], or HxWx3 uint8). Minimum size 8×8. A list of images gives a list of outputs. |
| Options | model(image, fit="pad", fov_x=None, apply_mask=True, force_projection=True). fit="stretch" resizes the image to the canvas. fov_x is a known horizontal field of view in degrees. from_pretrained(device_id=0, device=None, dispatch="eth", num_command_queues=None). The default is ETH dispatch, 1 CQ and a 12×10 compute grid (the p150 configuration). If the ETH open fails, the model gives a RuntimeWarning and uses worker dispatch. dispatch="worker", num_command_queues=2 is a Galaxy-only opt-in. |
Output (MoGeOutput) |
points float32 HxWx3 (metres), depth float32 HxW (metres), normal float32 HxWx3, mask bool HxW (True = valid), intrinsics (normalized 3×3), intrinsics_pixels (3×3), metric_scale, fov_x, fov_y. All maps have the input resolution. |
| Helpers | out.save_npz(path) writes the server's npz layout. out.to_dict() and out.to_torch() give the upstream infer() dict. out["depth"] also works. |
- Throughput:
model.predict_many(images)runs several images in order. The host loads the next image and post-processes the previous result while the device runs the current image. - Drop-in methods:
model.infer(image)has the same interface as the upstreamMoGeModel.infer.model.forward(image)is the raw forward on a 1920×1080 canvas. - Threads can share one model. The device runs one image at a time (batch 1). One model uses one chip.
- For the same image and options, the arrays are bit-identical to the
POST /predictresponse. The API givesnormalas float32. The server gives it as float16. - The model does not apply the EXIF orientation of JPEG files.
- Use the
code/of this repo for the Python API. The container image is older and does not containtt_moge.MoGeModel. - Full reference (all arguments, input forms, fields and speeds):
code/PYTHON.md. Runnable example:code/examples/quickstart.py.
Serving (HTTP)
tt-model pull changh95/moge-2-p150 --with-weights
tt-model serve changh95/moge-2-p150 # or with tt-cli: tt serve changh95/moge-2-p150
printf '{"image":"%s"}' "$(base64 -w0 media/source.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/moge-2-p150
- The weights go to your HF cache. The image does not contain them.
- The server uses port 20000 (or the next free port). It is ready when the log shows
Application startup complete. POST /predict:image(base64 PNG/JPEG); optionaloutput_format(npzdefault |png|json|npz_binary),fit(paddefault |stretch),fov_x(degrees),apply_mask(true),force_projection(true),include_depth_png(false).jsonis available only up to 512×512.GET /health,GET /info.
Response (shortened):
{"height": 1080, "width": 1920, "metric_scale": 10.34, "fov_x_deg": 86.40, "mask_coverage": 0.9055,
"intrinsics": [[0.5325, 0.0, 0.5], [0.0, 0.9466, 0.5], [0.0, 0.0, 1.0]],
"depth_m": {"min": 2.490, "median": 13.28, "max": 258.4},
"outputs": {"npz": "..."}, "timing_ms": {"device": 115.3, "total": 1342.1}}
outputs.npzis base64 ofnp.savez_compressedat the original resolution:pointsf32 [H,W,3] metres (OpenCV axes),depthf32 [H,W] metres,normalf16 [H,W,3],masku8 (1 = valid),intrinsicsf32 3×3 normalized (also returned asintrinsics_pixels),metric_scale. Invalid pixels haveinfdepth and points, and a zero normal.output_format: pngreturnsdepth_png16(metres = value ×encoding.depth_png_scale, 0 = invalid),normal_png(n = v/255·2−1) andmask_png(255 = valid).output_format: npz_binaryreturns the.npzfile as the response body (application/x-npz). The other JSON fields are in theX-MoGe-Metaheader. Only the currentcode/has this mode. The container image does not have it.
Demo
Demo & Performances
Warm, batch 1, one 1920×1080 image (1800 ViT tokens), bench_breakdown with 60 iterations per run (median / min). All rows use the p150 configuration: dispatch on ETH cores, 1 command queue (CQ), 12×10 compute grid. This is the default of the Python API and of the server. Measured on two Blackhole chips of a shared host on 2026-10-05 (chip 1 by the developer, chip 23 by an independent verifier).
| Metric | Performance |
|---|---|
Model call tt(image) (host input prep + H2D + trace + D2H + C++ host tail, 1920×1080 outputs) |
20.5 ms median (20.31–21.06, 10 runs on 2 chips) · 19.24 ms min |
| Device trace (DINOv2 ViT-L encoder + conv neck + heads) | 17.7–18.3 ms median · back-to-back loop 17.77–18.44 ms |
| First trace replay after idle (full clock), chip 1 | 16.01–16.18 ms |
| Host input prep · D2H (the head reads and the overlapped C++ host tail; with 1 CQ the reads run between the trace segments), chip 1 | 0.80–1.12 ms · 3.10–3.69 ms |
Python model(image) call, uint8 1920×1080 array (model call + focal and shift solve + finish at the input size), 4 processes on 2 chips |
25.1–25.5 ms median · 23.6–24.2 ms min (chip 1) · model.forward() 20.35–20.47 ms median · file path input 33.8–34.3 ms median |
Served /predict timing_ms.device, 1920×1080 PNG → npz, 30 requests, 5 runs on 2 chips |
19.4–20.3 ms median (19.2 min) · png 800×600 19.3–20.0 ms |
The ETH-dispatch configuration and the earlier worker-dispatch configuration (worker dispatch, 2 CQs) give bit-identical outputs. Accuracy against the torch fp32 reference on media/source.png: points / depth / normal / mask PCC 0.999953 / 0.999919 / 0.999958 / 0.999996. On 18 photos × 6 input draws, the mean aligned depth AbsRel is 0.00237 and the metric-scale error is 0.983 % RMS (max 2.78 %). model(image) adds about 5 ms of host post-processing to the model call. A file path also adds the image decode. The Python API changes no device code and no numerics: its arrays are identical to the HTTP server output (normal after the float16 cast). Details: VERIFICATION_2026-10-03.md, section "p150 ETH-dispatch compliance (2026-10-05)".
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the same 1920×1080 canvas. The "incl. H2D/D2H" column compares with our model call (20.5 ms). The "forward only" column compares with our device trace (17.8 ms). Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU incl. H2D/D2H | vs current build (20.5 ms) | GPU forward only vs ours (17.8 ms) |
|---|---|---|---|
| fp32 strict, eager | 65.5 ms | Blackhole 3.19× faster | 60.3 ms: Blackhole 3.39× faster |
| tf32, eager | 48.0 ms | Blackhole 2.34× faster | 42.3 ms: Blackhole 2.37× faster |
| bf16 autocast, eager | 30.9 ms | Blackhole 1.51× faster | 25.8 ms: Blackhole 1.45× faster |
| fp16 autocast, eager | 29.5 ms | Blackhole 1.44× faster | 24.2 ms: Blackhole 1.36× faster |
fp16 autocast + torch.compile (reduce-overhead) |
20.0 ms | GPU 1.02× faster | 14.9 ms: GPU 1.20× faster |
The Blackhole lead comes from one metal trace of fused, sharded bf16 kernels on all 120 compute cores. The trace runs in segments, one for each output head. With 1 CQ, the read of each output head is queued after its segment, and the C++ host tail of that head runs while the device computes the next heads. With fp16 and torch.compile (CUDA graphs), the GPU is about 1.2× faster on the forward alone and level on the model call. In the served device stage, the GPU (bf16) takes 31.7 ms and this build takes 19.7 ms (median of 5 runs).
Caveats
- Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (
patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication. - Worker (Tensix) dispatch with 2 CQs (
MOGE_DISPATCH=worker MOGE_2CQ=1, ordispatch="worker", num_command_queues=2in Python) is an opt-in for Galaxy chips only. On a p150, worker dispatch gives only an 11×10 compute grid, so the model is slower. The outputs are bit-identical. - Every image is placed on a fixed 1920×1080 canvas (1800 ViT tokens, 32×57 grid):
fit: padletterboxes and crops the border back out,stretchsquashes; one image per request, batch 1, requests are serialised on the chip. - The device uses a bf16 encoder and BFP8 conv weights, so the outputs differ slightly from the fp32 reference. Some optimizations also change numerics (HiFi2 matmuls with gain compensation, polynomial GELU, fused add + LayerNorm, HiFi3 ConvT as matmul).
OPT_REPORT.mdlists each change. Depth and points stay float32, because theexpremap reaches about 1e11 in invalid regions. - At boot, the server captures the device graph (ViT-L encoder, projection fold, conv neck, heads) as metal traces: one segment for each output head, plus the scale-head trace.
TT_FUSED=0restores the old eager-encoder path (about 215 ms device time in the first release). - Dense outputs are base64
npz(default) or 16-bit/8-bit PNGs inside the JSON envelope;jsonnested lists are refused above 512×512. - Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - The build uses tt-metal
v0.78.0-dev20260820(main8b98410e730). - p150a power was not measured, so no efficiency comparison is made.
Licensing
- Weights: Ruicheng/moge-2-vitl-normal, MIT (licence).
- Port and serving code (
code/): Apache-2.0, from changh95/tt-MoGe; the vendoredmicrosoft/MoGereference and utils3d undercode/tt_moge/reference/keep their MIT licence.
Provenance
These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build, see OPT_REPORT.md; 2026-10-04 Python API tt_moge.MoGeModel, see code/PYTHON.md; 2026-10-05 default of ETH dispatch with 1 CQ, see OPT_REPORT.md) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
36f1eb12a72865ea (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-13T15:29:57+00:00 by tt-model 0.1.0 |


