--- license: other license_name: mit-and-creativeml-openrail-m license_link: LICENSE.md base_model: timbrooks/instruct-pix2pix pipeline_tag: image-to-image tags: - core-ml - instruct-pix2pix - image-to-image - ios --- # InstructPix2Pix — Core ML (6-bit, for iPhone) This is a **converted copy** of [timbrooks/instruct-pix2pix](https://huggingface.co/timbrooks/instruct-pix2pix), created by **Tim Brooks, Aleksander Holynski and Alexei A. Efros** (UC Berkeley). It edits a photo from a written instruction, such as "make it winter" or "turn the car red". CommuniDev LLC did **not** train or fine-tune this model. We only changed its file format so it can run on an iPhone. It is published so that our iOS app [CommuniMuse](https://communidev.com) can download it on first use. ## What we changed We converted the model with Apple's [ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion) converter (`torch2coreml`, commit `ea2805dc`). There was no retraining. - Core ML, attention `SPLIT_EINSUM_V2` (Neural Engine layout), compute units `CPU_AND_NE`, 512 × 512 images. - UNet: - batch size 1, 8 input channels; - **6-bit palettized weights**; - split into two chunks (`UnetChunk1`, `UnetChunk2`) so it fits iPhone memory. - Text encoder: 6-bit palettized. - VAE encoder, VAE decoder and safety checker: fp16. - The weights are otherwise unchanged. Modified files: every `.mlmodelc` in `compiled/` is a converted, modified form of the original checkpoint. ## Files | Path | Size | |---|---| | `compiled/UnetChunk1.mlmodelc` | 319 MB | | `compiled/UnetChunk2.mlmodelc` | 300 MB | | `compiled/TextEncoder.mlmodelc` | 134 MB | | `compiled/VAEEncoder.mlmodelc` | 65 MB | | `compiled/VAEDecoder.mlmodelc` | 95 MB | | `compiled/SafetyChecker.mlmodelc` | 580 MB | | `compiled/vocab.json`, `compiled/merges.txt` | 1.3 MB | `SHA256SUMS` lists a SHA-256 hash for every file. Check it with `shasum -a 256 -c SHA256SUMS`. ## How it is used The model needs InstructPix2Pix's three-way classifier-free guidance. That is not the standard `StableDiffusionPipeline`: - text embeddings for `[prompt, negative, negative]`; - image latents `[image, image, zeros]`, where the image latents are the VAE mean, **not** scaled by 0.18215; - noise = uncond + g_text·(text − image) + g_img·(image − uncond). Defaults are g_text 7.5 and g_img 1.5, with DPM-Solver++ at 20–25 steps. The safety checker must stay on. ## Licence Two licences apply. Both are in `LICENSE.md`. 1. **MIT**: the InstructPix2Pix model and code, Copyright 2023 Timothy Brooks, Aleksander Holynski, Alexei A. Efros. 2. **CreativeML OpenRAIL-M**: InstructPix2Pix is fine-tuned from Stable Diffusion v1.5, so that licence's terms and **use-based restrictions** apply to this model too. **By downloading or using this model you agree to the CreativeML OpenRAIL-M use restrictions.** Among other things, you may not use it: - to break the law; - to exploit or harm minors; - to generate verifiably false information intended to harm others; - to defame, disparage or harass people; - to discriminate. The full list is in Attachment A of the licence. ## Limitations - Output is 512 × 512. Our app fits the photo into the square and crops the result back out. - Some instructions work weakly. Weather and season changes on night scenes are an example. - Portrait edits occasionally distort faces; retrying with another seed usually helps. - These are properties of the original model. ## Citation ```bibtex @article{brooks2022instructpix2pix, title = {InstructPix2Pix: Learning to Follow Image Editing Instructions}, author = {Brooks, Tim and Holynski, Aleksander and Efros, Alexei A.}, journal = {arXiv preprint arXiv:2211.09800}, year = {2022} } ```