Mirayu models
ONNX and GGUF models used by Mirayu, a free,
open-source desktop app for batch portrait retouching. The app downloads
them from here (and from their original sources) and checks every file
against the sha256 in its catalog
(catalog/models/<id>/manifest.toml).
None of these models were trained by the Mirayu project. Each folder is either an unchanged copy of a file published by its authors, or a format conversion (TFLite to ONNX, or a file taken out of a zip archive) with the steps to reproduce it below. Each folder carries the licence of its original source. The licence of each model is the licence of its original source, not of this repository.
| Folder | Model | Original source | Licence | Change made here |
|---|---|---|---|---|
mediapipe-face-landmarker/ |
MediaPipe Face Landmarker (face detector, 478-point mesh, blendshapes) | Google, face_landmarker.task float16 v1 |
Apache-2.0 | TFLite converted to ONNX (see below) |
rtmo-m-body7/ |
RTMO-m multi-person pose, Body7, COCO 17 keypoints | OpenMMLab MMPose, ONNX SDK zip | Apache-2.0 (see note) | end2end.onnx taken out of the zip, renamed |
yunet-2023mar/ |
YuNet face detector | OpenCV Zoo, Hugging Face opencv/face_detection_yunet @ 3cc26e7 |
MIT | none (mirror) |
sface-2021dec/ |
SFace face recognition | OpenCV Zoo, Hugging Face opencv/face_recognition_sface @ 3d70824 |
Apache-2.0 | none (mirror) |
bisenet-resnet18-celebamask/ |
BiSeNet ResNet-18 face parsing (CelebAMask-HQ, 19 classes) | yakhyo, face-parsing release weights (resnet18.onnx) |
code MIT; see note | renamed to bisenet_resnet18.onnx |
lama-carve-512/ |
LaMa (big-lama) inpainting, 512 x 512 | Carve, Hugging Face Carve/LaMa-ONNX @ c3c0c9e (lama_fp32.onnx), from advimman/lama |
Apache-2.0 | none (mirror) |
modnet-xenova/ |
MODNet portrait matting | Xenova, Hugging Face Xenova/modnet @ fa2fa54 (onnx/model.onnx), from ZHKKKe/MODNet |
Apache-2.0 | renamed to modnet.onnx |
Generative models (GGUF, for stable-diffusion.cpp)
Mirayu's optional generative fixes (teeth, hairline, glare and similar small masked repairs) run these with stable-diffusion.cpp. A diffusion model needs its text encoder and its VAE from the folders next to it:
- FLUX.2 klein 4B:
flux2-klein-4b/+qwen3-4b/+flux2-vae/ - Z-Image Turbo:
z-image-turbo/+qwen3-4b-instruct-2507/+flux1-ae/ - Qwen-Image-Edit-2511:
qwen-image-edit-2511/+qwen2.5-vl-7b-instruct/(both files) +qwen-image-vae/; optionally the 4-step LoRA inqwen-image-edit-2511-lightning/(4 steps, CFG 1)
| Folder | Model | Original source | Licence | Change made here |
|---|---|---|---|---|
flux2-klein-4b/ |
FLUX.2 [klein] 4B, GGUF Q4_K_M, Q5_K_M, Q8_0 | Black Forest Labs, black-forest-labs/FLUX.2-klein-4B @ e7b7dc2; GGUF: Q8_0 from leejet/FLUX.2-klein-4B-GGUF @ 3b1f5a9, Q4_K_M and Q5_K_M from unsloth/FLUX.2-klein-4B-GGUF @ 0084d1d |
Apache-2.0 | none (mirror) |
flux2-vae/ |
FLUX.2 VAE | black-forest-labs/FLUX.2-klein-4B @ e7b7dc2 (vae/diffusion_pytorch_model.safetensors) |
Apache-2.0 | renamed to flux2-vae.safetensors |
qwen3-4b/ |
Qwen3-4B text encoder, GGUF Q4_K_M, Q8_0 | Qwen, Qwen/Qwen3-4B @ 1cfa9a7; GGUF: unsloth/Qwen3-4B-GGUF @ 22c9fc8 |
Apache-2.0 | none (mirror) |
z-image-turbo/ |
Z-Image Turbo, GGUF Q8_0 | Tongyi-MAI, Tongyi-MAI/Z-Image-Turbo @ f332072; GGUF: leejet/Z-Image-Turbo-GGUF @ c61c0e4 |
Apache-2.0 | none (mirror) |
qwen3-4b-instruct-2507/ |
Qwen3-4B-Instruct-2507 text encoder, GGUF Q8_0 | Qwen, Qwen/Qwen3-4B-Instruct-2507 @ cdbee75; GGUF: unsloth/Qwen3-4B-Instruct-2507-GGUF @ a06e946 |
Apache-2.0 | none (mirror) |
flux1-ae/ |
FLUX.1 autoencoder (the VAE of FLUX.1 [schnell], also used by Z-Image) | Comfy-Org/z_image_turbo @ 6fc90a3 (split_files/vae/ae.safetensors), from Black Forest Labs' FLUX.1 [schnell] |
Apache-2.0 | renamed to flux1-ae.safetensors |
qwen-image-edit-2511/ |
Qwen-Image-Edit-2511, GGUF Q4_K_M, Q8_0 | Qwen, Qwen/Qwen-Image-Edit-2511 @ 6f3ccc0; GGUF: unsloth/Qwen-Image-Edit-2511-GGUF @ 0d33d96 |
Apache-2.0 | none (mirror) |
qwen2.5-vl-7b-instruct/ |
Qwen2.5-VL-7B-Instruct text encoder, GGUF Q8_0, and its vision tower (mmproj F16) | Qwen, Qwen/Qwen2.5-VL-7B-Instruct @ cc59489; GGUF: unsloth/Qwen2.5-VL-7B-Instruct-GGUF @ 68bb8bc |
Apache-2.0 | mmproj-F16.gguf renamed to Qwen2.5-VL-7B-Instruct-mmproj-F16.gguf |
qwen-image-vae/ |
Qwen-Image VAE | Comfy-Org/Qwen-Image_ComfyUI @ 1f12b17 (split_files/vae/qwen_image_vae.safetensors), from Qwen/Qwen-Image |
Apache-2.0 | none (mirror) |
qwen-image-edit-2511-lightning/ |
Qwen-Image-Edit-2511 Lightning, 4-step distillation LoRA (bf16) | LightX2V, lightx2v/Qwen-Image-Edit-2511-Lightning @ d74eba1 |
Apache-2.0 | none (mirror) |
The licence files are the ones published with each model: the FLUX.2
klein LICENSE.md, Qwen's LICENSE from each Qwen repository (the
Qwen-Image one for the Qwen-Image and Qwen2.5-VL folders, whose
repositories state Apache-2.0 without a licence file), the Z-Image GitHub
LICENSE (Tongyi-MAI/Z-Image @ 26f23ed),
model_licenses/LICENSE-FLUX1-schnell from black-forest-labs/flux
@ 802fb47 and the LightX2V GitHub LICENSE (ModelTC/LightX2V @ 8a97c75)
for the Lightning LoRA. All are Apache-2.0.
SHA256SUMS lists the hash of every model file.
Licence notes
- BiSeNet: the code and the ONNX export are MIT. The weights were trained on CelebAMask-HQ, whose images are for non-commercial research and educational use. Check whether that restricts your use of the model's outputs.
- RTMO Body7: code and weights are Apache-2.0. The Body7 training set combines COCO, AI Challenger, CrowdPose, MPII, sub-JHMDB, Halpe and PoseTrack18; some of those datasets are for research use only, so commercial use of the outputs is unclear.
- Models trained by InsightFace (ArcFace, SCRFD) are used by Mirayu but are not mirrored here: their licence is non-commercial research only and does not clearly allow redistribution. Mirayu downloads them from their publishers.
- FLUX.1 [dev] and FLUX.1 Fill [dev] are under the non-commercial FLUX.1 [dev] licence and gated on Hugging Face. They are not mirrored here: Mirayu downloads them with the user's own Hugging Face token after the user accepts the licence.
Attribution
- MediaPipe Face Landmarker, Google.
- RTMO: Peng Lu et al., "RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation", OpenMMLab MMPose.
- YuNet: Wei Wu, Hanyang Peng, Shiqi Yu, OpenCV Zoo.
- SFace: Yaoyao Zhong et al., OpenCV Zoo.
- BiSeNet face parsing: Yakhyokhuja Valikhujaev (yakhyo/face-parsing), after Changqian Yu et al. (BiSeNet) and Cheng-Han Lee et al. (CelebAMask-HQ).
- LaMa: Roman Suvorov et al. (Samsung AI Center); ONNX export by Carve.
- MODNet: Zhanghan Ke et al.; ONNX export by Xenova.
- FLUX.2 [klein] and FLUX.1 autoencoder: Black Forest Labs. GGUF conversions by leejet and Unsloth.
- Z-Image Turbo: Tongyi-MAI (Alibaba). GGUF conversion by leejet.
- Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen2.5-VL-7B-Instruct, Qwen-Image-Edit-2511, Qwen-Image VAE: Qwen team (Alibaba Cloud). GGUF conversions by Unsloth.
- Qwen-Image-Edit-2511 Lightning: LightX2V (ModelTC).
How the converted files were made
MediaPipe Face Landmarker (TFLite to ONNX)
The official task bundle is a zip with three TFLite models and the geometry metadata:
File in face_landmarker.task |
sha256 |
|---|---|
face_landmarker.task (the bundle) |
64184e229b263107bc2b804c6625db1341ff2bb731874b0bcc2fe6544e0bc9ff |
face_detector.tflite |
b4578f35940bf5a1a655214a1cce5cab13eba73c1297cd78e1a04c2380b0152f |
face_landmarks_detector.tflite |
c7d54204ce0448474c7f3fa9af494787c0965cbdd6f20fc72867e43046bd43d5 |
face_blendshapes.tflite |
4f36dded049db18d76048567439b2a7f58f1daabc00d78bfe8f3ad396a2d2082 |
geometry_pipeline_metadata_landmarks.binarypb (copied unchanged) |
bdbcda96dfcb7da883da124aaa2c55dee49770d934f0fcc71747f8c21bdc75b4 |
Each TFLite model was converted with tf2onnx at opset 17 (NHWC layout
kept; float16 weights become float32) by convert_mediapipe.py in that
folder (also in the Mirayu repository under prototype/scripts/). Tools:
Python 3.11.4, tensorflow-cpu 2.21.0, tf2onnx 1.17.0, onnx 1.23.1,
onnxruntime 1.30.0, numpy 2.4.6, protobuf 7.36.2.
Run from inside mediapipe-face-landmarker/:
python3.11 -m venv venv-tf
venv-tf/bin/pip install tensorflow-cpu==2.21.0 tf2onnx==1.17.0 onnx==1.23.1 onnxruntime==1.30.0
curl -LO https://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/1/face_landmarker.task
venv-tf/bin/python convert_mediapipe.py face_landmarker.task out --reference .
The script prints, for each model, the largest difference between TFLite and ONNX Runtime on random inputs (1.7e-4 for the detector, 4.1e-4 for the landmarks in 256 px units, 1.1e-6 for the blendshapes).
tf2onnx does not write byte-identical files from run to run (its generated
constant names and constant-folding order vary), so a fresh conversion has
other sha256 values than the files here. --reference . checks instead
that the fresh conversion gives exactly the same outputs as these files on
random inputs (maximum difference 0).
RTMO-m (zip member)
curl -LO https://download.openmmlab.com/mmpose/v1/projects/rtmo/onnx_sdk/rtmo-m_16xb16-600e_body7-640x640-39e78cc4_20231211.zip
# zip sha256 6b9d3be1323cc030444f731bbe593ea61f59b9481e23da2f8b60cb4b72136e16
unzip rtmo-m_16xb16-600e_body7-640x640-39e78cc4_20231211.zip end2end.onnx deploy.json detail.json pipeline.json
mv end2end.onnx rtmo_m_body7.onnx
The three JSON files from the zip are kept next to the model: they
describe its pre- and post-processing (input 1 x 3 x 640 x 640, BGR 0..255,
image letterboxed with grey 114; outputs dets N x 5 and keypoints
N x 17 x 3).
Inputs and outputs
How Mirayu calls each model (the Rust adapters in
crates/mirayu-models
and the Python reference in
prototype/mirayu_proto
have the details):
- YuNet: BGR 0..255, any size (Mirayu patches the fixed 640 x 640 input to free height and width in memory).
- SFace: 112 x 112 aligned face crop.
- MediaPipe Face Landmarker: detector 128 x 128 RGB 0..1; landmarks 256 x 256 RGB 0..1 (478 x 3 points and presence); blendshapes 146 x 2 landmarks to 52 scores.
- BiSeNet: 512 x 512 RGB, ImageNet mean / std; 19-class logits.
- RTMO: see above.
- MODNet: RGB normalised to -1..1, short side 512, sides multiples of 32; alpha 0..1.
- LaMa:
image1 x 3 x 512 x 512 RGB 0..1 andmask1 x 1 x 512 x 512 (1 = fill); output 0..255.
- Downloads last month
- -
4-bit
5-bit
8-bit