MoGe-2 ViT-L (normal) — unofficial Core ML conversion
This is an unofficial conversion. It is not made, endorsed or supported by the MoGe authors, Microsoft or Meta.
It is the network of Ruicheng/moge-2-vitl-normal turned
into a Core ML package for Apple silicon Macs, so the Re-Light Studio photo editor can run it without Python. The
weights are the original weights; only the file format and the numeric precision of some layers changed.
What is in this repo
| file | what |
|---|---|
moge2-vitl-normal-mixed.mlpackage.zip |
the Core ML package, zipped with ditto -c -k --keepParent (one top-level folder, moge2-vitl-normal-mixed.mlpackage) |
convert/moge_coreml.py |
the conversion script, as it was run |
LICENSE-MoGe-MIT.txt |
MoGe's licence (MIT, Copyright (c) Microsoft Corporation) |
LICENSE-DINOv2-Apache-2.0.txt |
the DINOv2 backbone's licence (Apache-2.0, Copyright (c) Meta Platforms, Inc. and affiliates) |
The zip is 658 540 801 bytes (773 MB unzipped). Check it before use:
SHA-256 8a1079ade8c5efd510876d4018991f9731ffcfdb2f32801325ff7525d72fc6fb
The first revision of this repo (11ac4c230596a0ed13c45bfdedca3b229a5f5ff8) holds the package with the landscape
and portrait functions only: 623 799 857 bytes, SHA-256
e0d3a37075721d1d7e8e0a25b2ef107f99457a9f44f5a1eb8fdf5d7a8445e464.
Source
- Code:
microsoft/MoGeat commit07444410f1e33f402353b99d6ccd26bd31e469e8. - Weights:
Ruicheng/moge-2-vitl-normalat revisioncb0e8bbd6b1e243589717c78e750b1ba4c093acf(331 M parameters). - Converted with
coremltools9.0 and PyTorch 2.14. The script carries a three-line shim for one cast that coremltools 9.0 folds wrongly under NumPy 2.5. It imports the two revisions above and a few test stills from the converter's own project, so it documents the conversion more than it re-runs on its own.
What the package does
One multifunction package with a function per common photo shape. The functions share one copy of the weights on disk: a function adds about 16 MB. Each takes an RGB image in 0..1, float32, already resized to the token grid, which is the grid MoGe takes for a photo of that shape at 3600 tokens.
| function | photo shape | input | tokens | output grid |
|---|---|---|---|---|
landscape |
4:3 | 1×3×728×966 | 52 × 69 | 832 × 1104 |
portrait |
3:4 | 1×3×966×728 | 69 × 52 | 1104 × 832 |
landscape_3x2 |
3:2 | 1×3×686×1022 | 49 × 73 | 784 × 1168 |
portrait_2x3 |
2:3 | 1×3×1022×686 | 73 × 49 | 1168 × 784 |
landscape_16x9 |
16:9 | 1×3×630×1120 | 45 × 80 | 720 × 1280 |
portrait_9x16 |
9:16 | 1×3×1120×630 | 80 × 45 | 1280 × 720 |
square |
1:1 | 1×3×840×840 | 60 × 60 | 960 × 960 |
Outputs (NCHW, float32): normal (unit vectors, OpenCV camera frame: x right, y down, z away), points (the
point map after MoGe's exp remap; channel 2 is the raw z), mask (after the sigmoid) and metric_scale.
This is the network only, as MoGeModel.forward runs it. MoGe's infer post-processing (focal length and shift
recovery, metric depth) is not in the package. Each function bakes in the position-embedding interpolation and the
UV planes for its own grid, so a photo of another shape needs a letterbox in the nearest function. A letterbox
is not free: black bars moved the normals of a face by 1.5 to 3° at the median in our tests, thin bars too.
In memory the weights are not shared: on an M2 Max a loaded function takes about 1.2 GB alone and 0.8 GB beside another, so load the functions you need and release the others.
Precision
Mixed: float16 through the network, float32 in the heads' last level at the full output grid. Plain float16
bands the depth (1 490 distinct z values in a face patch where float32 has 131 077); the mixed package keeps
130 819. Run it with MLComputeUnits.cpuAndGPU. On an M2 Max it takes 0.61 s a photo; the Neural Engine is slower
for this model.
Measured against the PyTorch float32 network on four stills cut to each function's shape (28 runs): normals within 0.22° at the 99th percentile, raw z within 1.7e-2 of its percentile-scaled range (within 1e-3 at the 99th percentile), masks identical, metric scale within 0.13 %.
Licences
- MoGe code and weights: MIT, Copyright (c) Microsoft Corporation. See
LICENSE-MoGe-MIT.txt. - DINOv2 backbone: Apache-2.0, Copyright (c) Meta Platforms, Inc. and affiliates. See
LICENSE-DINOv2-Apache-2.0.txt. The weights inside this package are a numeric conversion of the DINOv2-based encoder trained by the MoGe authors. - The conversion script and this card: MIT.
Citation
Please cite the MoGe-2 paper and the DINOv2 paper, not this conversion. See the MoGe repository for the entries.