Image-to-Image
core-ml
instruct-pix2pix
ios
lichtbizinc's picture
Core ML (6-bit, iPhone) conversion of timbrooks/instruct-pix2pix by CommuniDev LLC
34a0188 verified
|
Raw History Blame Contribute Delete
3.77 kB
---
license: other
license_name: mit-and-creativeml-openrail-m
license_link: LICENSE.md
base_model: timbrooks/instruct-pix2pix
pipeline_tag: image-to-image
tags:
- core-ml
- instruct-pix2pix
- image-to-image
- ios
---
# InstructPix2Pix — Core ML (6-bit, for iPhone)
This is a **converted copy** of [timbrooks/instruct-pix2pix](https://huggingface.co/timbrooks/instruct-pix2pix),
created by **Tim Brooks, Aleksander Holynski and Alexei A. Efros** (UC Berkeley). It edits a photo from a
written instruction, such as "make it winter" or "turn the car red".
CommuniDev LLC did **not** train or fine-tune this model. We only changed its file format so it can run
on an iPhone. It is published so that our iOS app [CommuniMuse](https://communidev.com) can download
it on first use.
## What we changed
We converted the model with Apple's [ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion)
converter (`torch2coreml`, commit `ea2805dc`). There was no retraining.
- Core ML, attention `SPLIT_EINSUM_V2` (Neural Engine layout), compute units `CPU_AND_NE`, 512 × 512 images.
- UNet:
- batch size 1, 8 input channels;
- **6-bit palettized weights**;
- split into two chunks (`UnetChunk1`, `UnetChunk2`) so it fits iPhone memory.
- Text encoder: 6-bit palettized.
- VAE encoder, VAE decoder and safety checker: fp16.
- The weights are otherwise unchanged.
Modified files: every `.mlmodelc` in `compiled/` is a converted, modified form of the original checkpoint.
## Files
| Path | Size |
|---|---|
| `compiled/UnetChunk1.mlmodelc` | 319 MB |
| `compiled/UnetChunk2.mlmodelc` | 300 MB |
| `compiled/TextEncoder.mlmodelc` | 134 MB |
| `compiled/VAEEncoder.mlmodelc` | 65 MB |
| `compiled/VAEDecoder.mlmodelc` | 95 MB |
| `compiled/SafetyChecker.mlmodelc` | 580 MB |
| `compiled/vocab.json`, `compiled/merges.txt` | 1.3 MB |
`SHA256SUMS` lists a SHA-256 hash for every file. Check it with `shasum -a 256 -c SHA256SUMS`.
## How it is used
The model needs InstructPix2Pix's three-way classifier-free guidance. That is not the standard
`StableDiffusionPipeline`:
- text embeddings for `[prompt, negative, negative]`;
- image latents `[image, image, zeros]`, where the image latents are the VAE mean, **not** scaled by 0.18215;
- noise = uncond + g_text·(text − image) + g_img·(image − uncond).
Defaults are g_text 7.5 and g_img 1.5, with DPM-Solver++ at 20–25 steps. The safety checker must stay on.
## Licence
Two licences apply. Both are in `LICENSE.md`.
1. **MIT**: the InstructPix2Pix model and code, Copyright 2023 Timothy Brooks, Aleksander Holynski, Alexei A. Efros.
2. **CreativeML OpenRAIL-M**: InstructPix2Pix is fine-tuned from Stable Diffusion v1.5, so that licence's
terms and **use-based restrictions** apply to this model too.
**By downloading or using this model you agree to the CreativeML OpenRAIL-M use restrictions.** Among
other things, you may not use it:
- to break the law;
- to exploit or harm minors;
- to generate verifiably false information intended to harm others;
- to defame, disparage or harass people;
- to discriminate.
The full list is in Attachment A of the licence.
## Limitations
- Output is 512 × 512. Our app fits the photo into the square and crops the result back out.
- Some instructions work weakly. Weather and season changes on night scenes are an example.
- Portrait edits occasionally distort faces; retrying with another seed usually helps.
- These are properties of the original model.
## Citation
```bibtex
@article{brooks2022instructpix2pix,
title = {InstructPix2Pix: Learning to Follow Image Editing Instructions},
author = {Brooks, Tim and Holynski, Aleksander and Efros, Alexei A.},
journal = {arXiv preprint arXiv:2211.09800},
year = {2022}
}
```