Image-to-Image
Diffusers
ONNX
Safetensors
StableDiffusionXLInpaintPipeline
stable-diffusion-xl
inpainting
virtual try-on
Instructions to use Devender113/IDM-VTON with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Devender113/IDM-VTON with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import AutoPipelineForInpainting from diffusers.utils import load_image # switch to "mps" for apple devices pipe = AutoPipelineForInpainting.from_pretrained("Devender113/IDM-VTON", dtype=torch.float16, device_map="cuda") img_url = "https://raw.githubusercontent.com/CompVis/latent-diffusion/main/data/inpainting_examples/overture-creations-5sI6fQgYIuo.png" mask_url = "https://raw.githubusercontent.com/CompVis/latent-diffusion/main/data/inpainting_examples/overture-creations-5sI6fQgYIuo_mask.png" image = load_image(img_url).resize((1024, 1024)) mask_image = load_image(mask_url).resize((1024, 1024)) prompt = "a tiger sitting on a park bench" generator = torch.Generator(device="cuda").manual_seed(0) image = pipe( prompt=prompt, image=image, mask_image=mask_image, guidance_scale=8.0, num_inference_steps=20, # steps between 15 and 30 work well for us strength=0.99, # make sure to use `strength` below 1.0 generator=generator, ).images[0] - Notebooks
- Google Colab
- Kaggle
Download image_encoder/config.json from Devender113/IDM-VTON: direct link, hf CLI and curl.
- Browser
- Download file 560 Bytes
-
https://huggingface.co/Devender113/IDM-VTON/resolve/main/image_encoder/config.json
- Command line
-
hf download hf://Devender113/IDM-VTON/image_encoder/config.json
-
curl -L -o config.json https://huggingface.co/Devender113/IDM-VTON/resolve/main/image_encoder/config.json
560 Bytes
| { | |
| "_name_or_path": "./image_encoder", | |
| "architectures": [ | |
| "CLIPVisionModelWithProjection" | |
| ], | |
| "attention_dropout": 0.0, | |
| "dropout": 0.0, | |
| "hidden_act": "gelu", | |
| "hidden_size": 1280, | |
| "image_size": 224, | |
| "initializer_factor": 1.0, | |
| "initializer_range": 0.02, | |
| "intermediate_size": 5120, | |
| "layer_norm_eps": 1e-05, | |
| "model_type": "clip_vision_model", | |
| "num_attention_heads": 16, | |
| "num_channels": 3, | |
| "num_hidden_layers": 32, | |
| "patch_size": 14, | |
| "projection_dim": 1024, | |
| "torch_dtype": "float16", | |
| "transformers_version": "4.28.0.dev0" | |
| } | |