MoGe-2 ViT-L (normal) — unofficial Core ML conversion

This is an unofficial conversion. It is not made, endorsed or supported by the MoGe authors, Microsoft or Meta. It is the network of Ruicheng/moge-2-vitl-normal turned into a Core ML package for Apple silicon Macs, so the Re-Light Studio photo editor can run it without Python. The weights are the original weights; only the file format and the numeric precision of some layers changed.

What is in this repo

file what
moge2-vitl-normal-mixed.mlpackage.zip the Core ML package, zipped with ditto -c -k --keepParent (one top-level folder, moge2-vitl-normal-mixed.mlpackage)
convert/moge_coreml.py the conversion script, as it was run
LICENSE-MoGe-MIT.txt MoGe's licence (MIT, Copyright (c) Microsoft Corporation)
LICENSE-DINOv2-Apache-2.0.txt the DINOv2 backbone's licence (Apache-2.0, Copyright (c) Meta Platforms, Inc. and affiliates)

The zip is 658 540 801 bytes (773 MB unzipped). Check it before use:

SHA-256  8a1079ade8c5efd510876d4018991f9731ffcfdb2f32801325ff7525d72fc6fb

The first revision of this repo (11ac4c230596a0ed13c45bfdedca3b229a5f5ff8) holds the package with the landscape and portrait functions only: 623 799 857 bytes, SHA-256 e0d3a37075721d1d7e8e0a25b2ef107f99457a9f44f5a1eb8fdf5d7a8445e464.

Source

  • Code: microsoft/MoGe at commit 07444410f1e33f402353b99d6ccd26bd31e469e8.
  • Weights: Ruicheng/moge-2-vitl-normal at revision cb0e8bbd6b1e243589717c78e750b1ba4c093acf (331 M parameters).
  • Converted with coremltools 9.0 and PyTorch 2.14. The script carries a three-line shim for one cast that coremltools 9.0 folds wrongly under NumPy 2.5. It imports the two revisions above and a few test stills from the converter's own project, so it documents the conversion more than it re-runs on its own.

What the package does

One multifunction package with a function per common photo shape. The functions share one copy of the weights on disk: a function adds about 16 MB. Each takes an RGB image in 0..1, float32, already resized to the token grid, which is the grid MoGe takes for a photo of that shape at 3600 tokens.

function photo shape input tokens output grid
landscape 4:3 1×3×728×966 52 × 69 832 × 1104
portrait 3:4 1×3×966×728 69 × 52 1104 × 832
landscape_3x2 3:2 1×3×686×1022 49 × 73 784 × 1168
portrait_2x3 2:3 1×3×1022×686 73 × 49 1168 × 784
landscape_16x9 16:9 1×3×630×1120 45 × 80 720 × 1280
portrait_9x16 9:16 1×3×1120×630 80 × 45 1280 × 720
square 1:1 1×3×840×840 60 × 60 960 × 960

Outputs (NCHW, float32): normal (unit vectors, OpenCV camera frame: x right, y down, z away), points (the point map after MoGe's exp remap; channel 2 is the raw z), mask (after the sigmoid) and metric_scale.

This is the network only, as MoGeModel.forward runs it. MoGe's infer post-processing (focal length and shift recovery, metric depth) is not in the package. Each function bakes in the position-embedding interpolation and the UV planes for its own grid, so a photo of another shape needs a letterbox in the nearest function. A letterbox is not free: black bars moved the normals of a face by 1.5 to 3° at the median in our tests, thin bars too.

In memory the weights are not shared: on an M2 Max a loaded function takes about 1.2 GB alone and 0.8 GB beside another, so load the functions you need and release the others.

Precision

Mixed: float16 through the network, float32 in the heads' last level at the full output grid. Plain float16 bands the depth (1 490 distinct z values in a face patch where float32 has 131 077); the mixed package keeps 130 819. Run it with MLComputeUnits.cpuAndGPU. On an M2 Max it takes 0.61 s a photo; the Neural Engine is slower for this model.

Measured against the PyTorch float32 network on four stills cut to each function's shape (28 runs): normals within 0.22° at the 99th percentile, raw z within 1.7e-2 of its percentile-scaled range (within 1e-3 at the 99th percentile), masks identical, metric scale within 0.13 %.

Licences

  • MoGe code and weights: MIT, Copyright (c) Microsoft Corporation. See LICENSE-MoGe-MIT.txt.
  • DINOv2 backbone: Apache-2.0, Copyright (c) Meta Platforms, Inc. and affiliates. See LICENSE-DINOv2-Apache-2.0.txt. The weights inside this package are a numeric conversion of the DINOv2-based encoder trained by the MoGe authors.
  • The conversion script and this card: MIT.

Citation

Please cite the MoGe-2 paper and the DINOv2 paper, not this conversion. See the MoGe repository for the entries.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kaziko/moge-2-vitl-normal-coreml

Finetuned
(5)
this model