Mirayu models

ONNX and GGUF models used by Mirayu, a free, open-source desktop app for batch portrait retouching. The app downloads them from here (and from their original sources) and checks every file against the sha256 in its catalog (catalog/models/<id>/manifest.toml).

None of these models were trained by the Mirayu project. Each folder is either an unchanged copy of a file published by its authors, or a format conversion (TFLite to ONNX, or a file taken out of a zip archive) with the steps to reproduce it below. Each folder carries the licence of its original source. The licence of each model is the licence of its original source, not of this repository.

Folder Model Original source Licence Change made here
mediapipe-face-landmarker/ MediaPipe Face Landmarker (face detector, 478-point mesh, blendshapes) Google, face_landmarker.task float16 v1 Apache-2.0 TFLite converted to ONNX (see below)
rtmo-m-body7/ RTMO-m multi-person pose, Body7, COCO 17 keypoints OpenMMLab MMPose, ONNX SDK zip Apache-2.0 (see note) end2end.onnx taken out of the zip, renamed
yunet-2023mar/ YuNet face detector OpenCV Zoo, Hugging Face opencv/face_detection_yunet @ 3cc26e7 MIT none (mirror)
sface-2021dec/ SFace face recognition OpenCV Zoo, Hugging Face opencv/face_recognition_sface @ 3d70824 Apache-2.0 none (mirror)
bisenet-resnet18-celebamask/ BiSeNet ResNet-18 face parsing (CelebAMask-HQ, 19 classes) yakhyo, face-parsing release weights (resnet18.onnx) code MIT; see note renamed to bisenet_resnet18.onnx
lama-carve-512/ LaMa (big-lama) inpainting, 512 x 512 Carve, Hugging Face Carve/LaMa-ONNX @ c3c0c9e (lama_fp32.onnx), from advimman/lama Apache-2.0 none (mirror)
modnet-xenova/ MODNet portrait matting Xenova, Hugging Face Xenova/modnet @ fa2fa54 (onnx/model.onnx), from ZHKKKe/MODNet Apache-2.0 renamed to modnet.onnx

Generative models (GGUF, for stable-diffusion.cpp)

Mirayu's optional generative fixes (teeth, hairline, glare and similar small masked repairs) run these with stable-diffusion.cpp. A diffusion model needs its text encoder and its VAE from the folders next to it:

  • FLUX.2 klein 4B: flux2-klein-4b/ + qwen3-4b/ + flux2-vae/
  • Z-Image Turbo: z-image-turbo/ + qwen3-4b-instruct-2507/ + flux1-ae/
  • Qwen-Image-Edit-2511: qwen-image-edit-2511/ + qwen2.5-vl-7b-instruct/ (both files) + qwen-image-vae/; optionally the 4-step LoRA in qwen-image-edit-2511-lightning/ (4 steps, CFG 1)
Folder Model Original source Licence Change made here
flux2-klein-4b/ FLUX.2 [klein] 4B, GGUF Q4_K_M, Q5_K_M, Q8_0 Black Forest Labs, black-forest-labs/FLUX.2-klein-4B @ e7b7dc2; GGUF: Q8_0 from leejet/FLUX.2-klein-4B-GGUF @ 3b1f5a9, Q4_K_M and Q5_K_M from unsloth/FLUX.2-klein-4B-GGUF @ 0084d1d Apache-2.0 none (mirror)
flux2-vae/ FLUX.2 VAE black-forest-labs/FLUX.2-klein-4B @ e7b7dc2 (vae/diffusion_pytorch_model.safetensors) Apache-2.0 renamed to flux2-vae.safetensors
qwen3-4b/ Qwen3-4B text encoder, GGUF Q4_K_M, Q8_0 Qwen, Qwen/Qwen3-4B @ 1cfa9a7; GGUF: unsloth/Qwen3-4B-GGUF @ 22c9fc8 Apache-2.0 none (mirror)
z-image-turbo/ Z-Image Turbo, GGUF Q8_0 Tongyi-MAI, Tongyi-MAI/Z-Image-Turbo @ f332072; GGUF: leejet/Z-Image-Turbo-GGUF @ c61c0e4 Apache-2.0 none (mirror)
qwen3-4b-instruct-2507/ Qwen3-4B-Instruct-2507 text encoder, GGUF Q8_0 Qwen, Qwen/Qwen3-4B-Instruct-2507 @ cdbee75; GGUF: unsloth/Qwen3-4B-Instruct-2507-GGUF @ a06e946 Apache-2.0 none (mirror)
flux1-ae/ FLUX.1 autoencoder (the VAE of FLUX.1 [schnell], also used by Z-Image) Comfy-Org/z_image_turbo @ 6fc90a3 (split_files/vae/ae.safetensors), from Black Forest Labs' FLUX.1 [schnell] Apache-2.0 renamed to flux1-ae.safetensors
qwen-image-edit-2511/ Qwen-Image-Edit-2511, GGUF Q4_K_M, Q8_0 Qwen, Qwen/Qwen-Image-Edit-2511 @ 6f3ccc0; GGUF: unsloth/Qwen-Image-Edit-2511-GGUF @ 0d33d96 Apache-2.0 none (mirror)
qwen2.5-vl-7b-instruct/ Qwen2.5-VL-7B-Instruct text encoder, GGUF Q8_0, and its vision tower (mmproj F16) Qwen, Qwen/Qwen2.5-VL-7B-Instruct @ cc59489; GGUF: unsloth/Qwen2.5-VL-7B-Instruct-GGUF @ 68bb8bc Apache-2.0 mmproj-F16.gguf renamed to Qwen2.5-VL-7B-Instruct-mmproj-F16.gguf
qwen-image-vae/ Qwen-Image VAE Comfy-Org/Qwen-Image_ComfyUI @ 1f12b17 (split_files/vae/qwen_image_vae.safetensors), from Qwen/Qwen-Image Apache-2.0 none (mirror)
qwen-image-edit-2511-lightning/ Qwen-Image-Edit-2511 Lightning, 4-step distillation LoRA (bf16) LightX2V, lightx2v/Qwen-Image-Edit-2511-Lightning @ d74eba1 Apache-2.0 none (mirror)

The licence files are the ones published with each model: the FLUX.2 klein LICENSE.md, Qwen's LICENSE from each Qwen repository (the Qwen-Image one for the Qwen-Image and Qwen2.5-VL folders, whose repositories state Apache-2.0 without a licence file), the Z-Image GitHub LICENSE (Tongyi-MAI/Z-Image @ 26f23ed), model_licenses/LICENSE-FLUX1-schnell from black-forest-labs/flux @ 802fb47 and the LightX2V GitHub LICENSE (ModelTC/LightX2V @ 8a97c75) for the Lightning LoRA. All are Apache-2.0.

SHA256SUMS lists the hash of every model file.

Licence notes

  • BiSeNet: the code and the ONNX export are MIT. The weights were trained on CelebAMask-HQ, whose images are for non-commercial research and educational use. Check whether that restricts your use of the model's outputs.
  • RTMO Body7: code and weights are Apache-2.0. The Body7 training set combines COCO, AI Challenger, CrowdPose, MPII, sub-JHMDB, Halpe and PoseTrack18; some of those datasets are for research use only, so commercial use of the outputs is unclear.
  • Models trained by InsightFace (ArcFace, SCRFD) are used by Mirayu but are not mirrored here: their licence is non-commercial research only and does not clearly allow redistribution. Mirayu downloads them from their publishers.
  • FLUX.1 [dev] and FLUX.1 Fill [dev] are under the non-commercial FLUX.1 [dev] licence and gated on Hugging Face. They are not mirrored here: Mirayu downloads them with the user's own Hugging Face token after the user accepts the licence.

Attribution

  • MediaPipe Face Landmarker, Google.
  • RTMO: Peng Lu et al., "RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation", OpenMMLab MMPose.
  • YuNet: Wei Wu, Hanyang Peng, Shiqi Yu, OpenCV Zoo.
  • SFace: Yaoyao Zhong et al., OpenCV Zoo.
  • BiSeNet face parsing: Yakhyokhuja Valikhujaev (yakhyo/face-parsing), after Changqian Yu et al. (BiSeNet) and Cheng-Han Lee et al. (CelebAMask-HQ).
  • LaMa: Roman Suvorov et al. (Samsung AI Center); ONNX export by Carve.
  • MODNet: Zhanghan Ke et al.; ONNX export by Xenova.
  • FLUX.2 [klein] and FLUX.1 autoencoder: Black Forest Labs. GGUF conversions by leejet and Unsloth.
  • Z-Image Turbo: Tongyi-MAI (Alibaba). GGUF conversion by leejet.
  • Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen2.5-VL-7B-Instruct, Qwen-Image-Edit-2511, Qwen-Image VAE: Qwen team (Alibaba Cloud). GGUF conversions by Unsloth.
  • Qwen-Image-Edit-2511 Lightning: LightX2V (ModelTC).

How the converted files were made

MediaPipe Face Landmarker (TFLite to ONNX)

The official task bundle is a zip with three TFLite models and the geometry metadata:

File in face_landmarker.task sha256
face_landmarker.task (the bundle) 64184e229b263107bc2b804c6625db1341ff2bb731874b0bcc2fe6544e0bc9ff
face_detector.tflite b4578f35940bf5a1a655214a1cce5cab13eba73c1297cd78e1a04c2380b0152f
face_landmarks_detector.tflite c7d54204ce0448474c7f3fa9af494787c0965cbdd6f20fc72867e43046bd43d5
face_blendshapes.tflite 4f36dded049db18d76048567439b2a7f58f1daabc00d78bfe8f3ad396a2d2082
geometry_pipeline_metadata_landmarks.binarypb (copied unchanged) bdbcda96dfcb7da883da124aaa2c55dee49770d934f0fcc71747f8c21bdc75b4

Each TFLite model was converted with tf2onnx at opset 17 (NHWC layout kept; float16 weights become float32) by convert_mediapipe.py in that folder (also in the Mirayu repository under prototype/scripts/). Tools: Python 3.11.4, tensorflow-cpu 2.21.0, tf2onnx 1.17.0, onnx 1.23.1, onnxruntime 1.30.0, numpy 2.4.6, protobuf 7.36.2.

Run from inside mediapipe-face-landmarker/:

python3.11 -m venv venv-tf
venv-tf/bin/pip install tensorflow-cpu==2.21.0 tf2onnx==1.17.0 onnx==1.23.1 onnxruntime==1.30.0
curl -LO https://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/1/face_landmarker.task
venv-tf/bin/python convert_mediapipe.py face_landmarker.task out --reference .

The script prints, for each model, the largest difference between TFLite and ONNX Runtime on random inputs (1.7e-4 for the detector, 4.1e-4 for the landmarks in 256 px units, 1.1e-6 for the blendshapes).

tf2onnx does not write byte-identical files from run to run (its generated constant names and constant-folding order vary), so a fresh conversion has other sha256 values than the files here. --reference . checks instead that the fresh conversion gives exactly the same outputs as these files on random inputs (maximum difference 0).

RTMO-m (zip member)

curl -LO https://download.openmmlab.com/mmpose/v1/projects/rtmo/onnx_sdk/rtmo-m_16xb16-600e_body7-640x640-39e78cc4_20231211.zip
# zip sha256 6b9d3be1323cc030444f731bbe593ea61f59b9481e23da2f8b60cb4b72136e16
unzip rtmo-m_16xb16-600e_body7-640x640-39e78cc4_20231211.zip end2end.onnx deploy.json detail.json pipeline.json
mv end2end.onnx rtmo_m_body7.onnx

The three JSON files from the zip are kept next to the model: they describe its pre- and post-processing (input 1 x 3 x 640 x 640, BGR 0..255, image letterboxed with grey 114; outputs dets N x 5 and keypoints N x 17 x 3).

Inputs and outputs

How Mirayu calls each model (the Rust adapters in crates/mirayu-models and the Python reference in prototype/mirayu_proto have the details):

  • YuNet: BGR 0..255, any size (Mirayu patches the fixed 640 x 640 input to free height and width in memory).
  • SFace: 112 x 112 aligned face crop.
  • MediaPipe Face Landmarker: detector 128 x 128 RGB 0..1; landmarks 256 x 256 RGB 0..1 (478 x 3 points and presence); blendshapes 146 x 2 landmarks to 52 scores.
  • BiSeNet: 512 x 512 RGB, ImageNet mean / std; 19-class logits.
  • RTMO: see above.
  • MODNet: RGB normalised to -1..1, short side 512, sides multiples of 32; alpha 0..1.
  • LaMa: image 1 x 3 x 512 x 512 RGB 0..1 and mask 1 x 1 x 512 x 512 (1 = fill); output 0..255.
Downloads last month
-
GGUF
Model size
4B params
Architecture
flux
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support