moge-2-p150

Microsoft MoGe-2 (ViT-L/normal checkpoint) port on one Tenstorrent Blackhole p150a. Weights: Ruicheng/moge-2-vitl-normal · Paper: arXiv:2507.02546 · Upstream code: microsoft/MoGe · Port: changh95/tt-MoGe

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a tt-metal / ttnn environment at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.

hf download changh95/moge-2-p150 --exclude "image/*" --local-dir moge-2-p150 && cd moge-2-p150
pip install -e code/        # adds numpy<2, pillow, scipy, huggingface_hub, opencv-python-headless
pip install -e "code/[server,test]"     # optional: also the HTTP server (fastapi, uvicorn, pydantic, isal) and pytest

Run the commands from the repo root (the directory that contains code/ and media/). The relative paths in this section (code/..., media/...) start at the repo root.

from tt_moge import MoGeModel

with MoGeModel.from_pretrained() as model:      # opens device 0, gets the weights from the HF cache
    out = model("photo.jpg")                    # a path, PIL image, numpy array or torch tensor

print(out)                  # MoGeOutput(1920x1080, valid 90.5%, median depth 13.322 m, fov_x 86.1 deg, metric_scale 10.3329)
depth = out.depth           # (H, W) float32, metres; inf where out.mask is False
points = out.points         # (H, W, 3) float32, metres, camera space (x right, y down, z forward)
normal = out.normal         # (H, W, 3) float32, unit normals
K = out.intrinsics_pixels   # (3, 3) camera matrix in pixels
  • from_pretrained downloads the Ruicheng/moge-2-vitl-normal weights to your HF cache, compiles the kernels and captures the device trace. The first load takes about 40 s. A load with a warm kernel cache takes about 8 s.
  • from_pretrained() also warms up the common call types. Your first call is then as fast as later calls (about 25 ms for a 1920x1080 image). To warm up other input types or sizes, use warmup_variants= or model.warmup(...) (see PYTHON.md).
  • Load the model one time and call it many times. A warm call on a 1920×1080 image takes about 25 ms.
  • The with block releases the trace and closes the device at the end. Without with, call model.close().
  • python code/examples/quickstart.py runs this snippet on media/source.png. It writes depth.png, normal.png, result.npz and result.json to ./moge_output/.
Input (image) A file path, a PIL.Image, a numpy array (HxWx3 uint8, or float in [0, 1]; also 3xHxW and grayscale), or a torch tensor (3xHxW / 1x3xHxW float in [0, 1], or HxWx3 uint8). Minimum size 8×8. A list of images gives a list of outputs.
Options model(image, fit="pad", fov_x=None, apply_mask=True, force_projection=True). fit="stretch" resizes the image to the canvas. fov_x is a known horizontal field of view in degrees. from_pretrained(device_id=0, device=None, dispatch="eth", num_command_queues=None). The default is ETH dispatch, 1 CQ and a 12×10 compute grid (the p150 configuration). If the ETH open fails, the model gives a RuntimeWarning and uses worker dispatch. dispatch="worker", num_command_queues=2 is a Galaxy-only opt-in.
Output (MoGeOutput) points float32 HxWx3 (metres), depth float32 HxW (metres), normal float32 HxWx3, mask bool HxW (True = valid), intrinsics (normalized 3×3), intrinsics_pixels (3×3), metric_scale, fov_x, fov_y. All maps have the input resolution.
Helpers out.save_npz(path) writes the server's npz layout. out.to_dict() and out.to_torch() give the upstream infer() dict. out["depth"] also works.
  • Throughput: model.predict_many(images) runs several images in order. The host loads the next image and post-processes the previous result while the device runs the current image.
  • Drop-in methods: model.infer(image) has the same interface as the upstream MoGeModel.infer. model.forward(image) is the raw forward on a 1920×1080 canvas.
  • Threads can share one model. The device runs one image at a time (batch 1). One model uses one chip.
  • For the same image and options, the arrays are bit-identical to the POST /predict response. The API gives normal as float32. The server gives it as float16.
  • The model does not apply the EXIF orientation of JPEG files.
  • Use the code/ of this repo for the Python API. The container image is older and does not contain tt_moge.MoGeModel.
  • Full reference (all arguments, input forms, fields and speeds): code/PYTHON.md. Runnable example: code/examples/quickstart.py.

Serving (HTTP)

tt-model pull  changh95/moge-2-p150 --with-weights
tt-model serve changh95/moge-2-p150        # or with tt-cli: tt serve changh95/moge-2-p150
printf '{"image":"%s"}' "$(base64 -w0 media/source.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/moge-2-p150
  • The weights go to your HF cache. The image does not contain them.
  • The server uses port 20000 (or the next free port). It is ready when the log shows Application startup complete.
  • POST /predict: image (base64 PNG/JPEG); optional output_format (npz default | png | json | npz_binary), fit (pad default | stretch), fov_x (degrees), apply_mask (true), force_projection (true), include_depth_png (false). json is available only up to 512×512.
  • GET /health, GET /info.

Response (shortened):

{"height": 1080, "width": 1920, "metric_scale": 10.34, "fov_x_deg": 86.40, "mask_coverage": 0.9055,
 "intrinsics": [[0.5325, 0.0, 0.5], [0.0, 0.9466, 0.5], [0.0, 0.0, 1.0]],
 "depth_m": {"min": 2.490, "median": 13.28, "max": 258.4},
 "outputs": {"npz": "..."}, "timing_ms": {"device": 115.3, "total": 1342.1}}
  • outputs.npz is base64 of np.savez_compressed at the original resolution: points f32 [H,W,3] metres (OpenCV axes), depth f32 [H,W] metres, normal f16 [H,W,3], mask u8 (1 = valid), intrinsics f32 3×3 normalized (also returned as intrinsics_pixels), metric_scale. Invalid pixels have inf depth and points, and a zero normal.
  • output_format: png returns depth_png16 (metres = value × encoding.depth_png_scale, 0 = invalid), normal_png (n = v/255·2−1) and mask_png (255 = valid).
  • output_format: npz_binary returns the .npz file as the response body (application/x-npz). The other JSON fields are in the X-MoGe-Meta header. Only the current code/ has this mode. The container image does not have it.

Demo

Input (media/source.png) Depth on p150a (media/depth.png) Normals on p150a (media/normal.png)

Demo & Performances

Warm, batch 1, one 1920×1080 image (1800 ViT tokens), bench_breakdown with 60 iterations per run (median / min). All rows use the p150 configuration: dispatch on ETH cores, 1 command queue (CQ), 12×10 compute grid. This is the default of the Python API and of the server. Measured on two Blackhole chips of a shared host on 2026-10-05 (chip 1 by the developer, chip 23 by an independent verifier).

Metric Performance
Model call tt(image) (host input prep + H2D + trace + D2H + C++ host tail, 1920×1080 outputs) 20.5 ms median (20.31–21.06, 10 runs on 2 chips) · 19.24 ms min
Device trace (DINOv2 ViT-L encoder + conv neck + heads) 17.7–18.3 ms median · back-to-back loop 17.77–18.44 ms
First trace replay after idle (full clock), chip 1 16.01–16.18 ms
Host input prep · D2H (the head reads and the overlapped C++ host tail; with 1 CQ the reads run between the trace segments), chip 1 0.80–1.12 ms · 3.10–3.69 ms
Python model(image) call, uint8 1920×1080 array (model call + focal and shift solve + finish at the input size), 4 processes on 2 chips 25.1–25.5 ms median · 23.6–24.2 ms min (chip 1) · model.forward() 20.35–20.47 ms median · file path input 33.8–34.3 ms median
Served /predict timing_ms.device, 1920×1080 PNG → npz, 30 requests, 5 runs on 2 chips 19.4–20.3 ms median (19.2 min) · png 800×600 19.3–20.0 ms

The ETH-dispatch configuration and the earlier worker-dispatch configuration (worker dispatch, 2 CQs) give bit-identical outputs. Accuracy against the torch fp32 reference on media/source.png: points / depth / normal / mask PCC 0.999953 / 0.999919 / 0.999958 / 0.999996. On 18 photos × 6 input draws, the mean aligned depth AbsRel is 0.00237 and the metric-scale error is 0.983 % RMS (max 2.78 %). model(image) adds about 5 ms of host post-processing to the model call. A file path also adds the image decode. The Python API changes no device code and no numerics: its arrays are identical to the HTTP server output (normal after the float16 cast). Details: VERIFICATION_2026-10-03.md, section "p150 ETH-dispatch compliance (2026-10-05)".

RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the same 1920×1080 canvas. The "incl. H2D/D2H" column compares with our model call (20.5 ms). The "forward only" column compares with our device trace (17.8 ms). Full table: GPU_COMPARISON.md.

RTX 5090 precision GPU incl. H2D/D2H vs current build (20.5 ms) GPU forward only vs ours (17.8 ms)
fp32 strict, eager 65.5 ms Blackhole 3.19× faster 60.3 ms: Blackhole 3.39× faster
tf32, eager 48.0 ms Blackhole 2.34× faster 42.3 ms: Blackhole 2.37× faster
bf16 autocast, eager 30.9 ms Blackhole 1.51× faster 25.8 ms: Blackhole 1.45× faster
fp16 autocast, eager 29.5 ms Blackhole 1.44× faster 24.2 ms: Blackhole 1.36× faster
fp16 autocast + torch.compile (reduce-overhead) 20.0 ms GPU 1.02× faster 14.9 ms: GPU 1.20× faster

The Blackhole lead comes from one metal trace of fused, sharded bf16 kernels on all 120 compute cores. The trace runs in segments, one for each output head. With 1 CQ, the read of each output head is queued after its segment, and the C++ host tail of that head runs while the device computes the next heads. With fp16 and torch.compile (CUDA graphs), the GPU is about 1.2× faster on the forward alone and level on the model call. In the served device stage, the GPU (bf16) takes 31.7 ms and this build takes 19.7 ms (median of 5 runs).

Caveats

  • Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
  • Worker (Tensix) dispatch with 2 CQs (MOGE_DISPATCH=worker MOGE_2CQ=1, or dispatch="worker", num_command_queues=2 in Python) is an opt-in for Galaxy chips only. On a p150, worker dispatch gives only an 11×10 compute grid, so the model is slower. The outputs are bit-identical.
  • Every image is placed on a fixed 1920×1080 canvas (1800 ViT tokens, 32×57 grid): fit: pad letterboxes and crops the border back out, stretch squashes; one image per request, batch 1, requests are serialised on the chip.
  • The device uses a bf16 encoder and BFP8 conv weights, so the outputs differ slightly from the fp32 reference. Some optimizations also change numerics (HiFi2 matmuls with gain compensation, polynomial GELU, fused add + LayerNorm, HiFi3 ConvT as matmul). OPT_REPORT.md lists each change. Depth and points stay float32, because the exp remap reaches about 1e11 in invalid regions.
  • At boot, the server captures the device graph (ViT-L encoder, projection fold, conv neck, heads) as metal traces: one segment for each output head, plus the scale-head trace. TT_FUSED=0 restores the old eager-encoder path (about 215 ms device time in the first release).
  • Dense outputs are base64 npz (default) or 16-bit/8-bit PNGs inside the JSON envelope; json nested lists are refused above 512×512.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • The build uses tt-metal v0.78.0-dev20260820 (main 8b98410e730).
  • p150a power was not measured, so no efficiency comparison is made.

Licensing

Provenance

These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build, see OPT_REPORT.md; 2026-10-04 Python API tt_moge.MoGeModel, see code/PYTHON.md; 2026-10-05 default of ETH dispatch with 1 CQ, see OPT_REPORT.md) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:

component built from
tt-metal 8b98410e730bb504fea43a88609756e34821d91d
code/ digest (image) 36f1eb12a72865ea (sha256, first 16 hex digits; the current code/ differs)
built 2026-09-13T15:29:57+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/moge-2-p150

Finetuned
(5)
this model

Collection including changh95/moge-2-p150

Paper for changh95/moge-2-p150