Improving Diffusion Models for Virtual Try-on
Paper • 2403.05139 • Published • 7
How to use MnLgt/IDM-VTON with Diffusers:
pip install -U diffusers transformers accelerate
import torch
from diffusers import AutoPipelineForInpainting
from diffusers.utils import load_image
# switch to "mps" for apple devices
pipe = AutoPipelineForInpainting.from_pretrained("MnLgt/IDM-VTON", dtype=torch.float16, device_map="cuda")
img_url = "https://raw.githubusercontent.com/CompVis/latent-diffusion/main/data/inpainting_examples/overture-creations-5sI6fQgYIuo.png"
mask_url = "https://raw.githubusercontent.com/CompVis/latent-diffusion/main/data/inpainting_examples/overture-creations-5sI6fQgYIuo_mask.png"
image = load_image(img_url).resize((1024, 1024))
mask_image = load_image(mask_url).resize((1024, 1024))
prompt = "a tiger sitting on a park bench"
generator = torch.Generator(device="cuda").manual_seed(0)
image = pipe(
prompt=prompt,
image=image,
mask_image=mask_image,
guidance_scale=8.0,
num_inference_steps=20, # steps between 15 and 30 work well for us
strength=0.99, # make sure to use `strength` below 1.0
generator=generator,
).images[0]This is an official implementation of paper 'Improving Diffusion Models for Authentic Virtual Try-on in the Wild'
🤗 Try our huggingface Demo
For the demo, GPUs are supported from zerogpu, and auto masking generation codes are based on OOTDiffusion and DCI-VTON.
Parts of the code are based on IP-Adapter.
@article{choi2024improving,
title={Improving Diffusion Models for Virtual Try-on},
author={Choi, Yisol and Kwak, Sangkyung and Lee, Kyungmin and Choi, Hyungwon and Shin, Jinwoo},
journal={arXiv preprint arXiv:2403.05139},
year={2024}
}
The codes and checkpoints in this repository are under the CC BY-NC-SA 4.0 license.