How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-Edit-2509", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("ArtmeScienceLab/Garments2Look-LoRA")

prompt = "Turn this cat into a dog"
input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png")

image = pipe(image=input_image, prompt=prompt).images[0]

Garments2Look-LoRA

Task-specific LoRA adapters for Qwen-Image-Edit-2509, trained for outfit-level virtual try-on with multiple garments and accessories.

Release status: both task-specific LoRA checkpoints, training/inference code, and updated dataset inputs are publicly available. Use the code repository for installation and inference, and the dataset card for download and preparation.

The released adapters were trained further than the CVPR rebuttal-stage models. With additional training and more suitable inpainting masks, we expect improved results compared with the earlier models.

The updated dataset provides 98,012 outfit records with v1.1 annotations, OOTD collages, editing source images, and five annotation types: ATR, DensePose, DWPose, LIP, and refined v3 masks with dilation. For inpainting, prepare Figure 1 using the v3 mask; for editing, use the provided edited/banana/ source image. Dataset preparation instructions and split counts are maintained in the dataset card.

Checkpoints

Task Checkpoint Training
Inpainting Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors 20K samples, 2 completed epochs, LoRA rank 32
Editing Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors 20K samples, 2 completed epochs, LoRA rank 32

These files are LoRA adapters, not complete base models. epoch-1 is zero-indexed and denotes the checkpoint saved after the second epoch. Use the checkpoint matching your task.

Garments2Look-LoRA/
β”œβ”€β”€ README.md
β”œβ”€β”€ examples/comparison.jpg
β”œβ”€β”€ manifest.json
β”œβ”€β”€ Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors
└── Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors

Inputs

Both tasks take two images and a text prompt:

Input Inpainting Editing
Figure 1 Person image with clothing regions masked in gray (128) Source person image wearing an existing outfit
Figure 2 OOTD collage of target reference items OOTD collage of target reference items
Prompt Numbered items, styling instructions, and layering order Numbered items, styling instructions, and layering order

Keep the collage item order consistent with the numbered prompt. The output is a person image wearing the target outfit. No separate mask argument is passed to inference: inpainting uses the already masked person image as Figure 1.

Installation and download

Run from the code repository root. The tested environment uses Python 3.10, PyTorch 2.7.1 with CUDA 12.8, and an NVIDIA H200.

git clone https://github.com/ArtmeScienceLab/Garments2Look.git
cd Garments2Look
conda create -n g2l-lora python=3.10 -y
conda activate g2l-lora
python -m pip install torch==2.7.1 torchvision==0.22.1 --index-url https://download.pytorch.org/whl/cu128
python -m pip install -r requirements.txt
python -m pip install -e . --no-deps

hf download Qwen/Qwen-Image-Edit-2509 --local-dir models/Qwen-Image-Edit-2509

# Download the released task-specific adapters.
hf download ArtmeScienceLab/Garments2Look-LoRA --local-dir models/Garments2Look-LoRA

Inpainting inference

python scripts/inference/inference.py --task inpainting \
  --model-dir models/Qwen-Image-Edit-2509 \
  --lora models/Garments2Look-LoRA/Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors \
  --origin examples/input-inpainting.png --ootd examples/ootd.png \
  --prompt-file examples/prompt-inpainting.txt \
  --output output/inpainting.png --seed 123 --steps 40 --cfg-scale 4.0

Editing inference

python scripts/inference/inference.py --task editing \
  --model-dir models/Qwen-Image-Edit-2509 \
  --lora models/Garments2Look-LoRA/Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors \
  --origin examples/input-editing.png --ootd examples/ootd.png \
  --prompt-file examples/prompt-editing.txt \
  --output output/editing.png --seed 123 --steps 40 --cfg-scale 4.0

The default inference seed is 123. Each run saves an output PNG and a JSON record containing its prompt and parameters. Images are aligned to multiples of 16 within a 1,048,576-pixel budget. Set CUDA_VISIBLE_DEVICES to select a GPU. For a custom outfit, replace the origin image, OOTD collage, and prompt file.

Use the existing server weights

The same commands work with absolute local paths; no adapter download is needed:

# Run from /home/hjy/repo/ours/Garments2Look/github-repo.
export BASE_MODEL=/mount/data/hjy/models/Qwen/Qwen-Image-Edit-2509
export LORA_DIR=/mount/data/hjy/models/Garments2Look-LoRA

python scripts/inference/inference.py --task inpainting \
  --model-dir "$BASE_MODEL" --lora "$LORA_DIR/Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors" \
  --origin examples/input-inpainting.png --ootd examples/ootd.png \
  --prompt-file examples/prompt-inpainting.txt --output output/inpainting.png

python scripts/inference/inference.py --task editing \
  --model-dir "$BASE_MODEL" --lora "$LORA_DIR/Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors" \
  --origin examples/input-editing.png --ootd examples/ootd.png \
  --prompt-file examples/prompt-editing.txt --output output/editing.png

Example

Inpainting and editing, seed 123

Columns: OOTD, inpainting input/output, editing input/output. The full prompts appear as two lines below the images. This six-item test example uses seed 123, 40 steps, and guidance 4.0 for both tasks.

Full prompt:

Keep the woman's identity, pose, background in Figure 1 unchanged, wearing the outfit in Figure 2, include (1) a top (partially unbuttoned, tucked-in), (2) a sweater (unbuttoned), (3) pants, (4) loafers, (5) a bag, (6) a belt (worn around waist). Layering Order: (1) -> (6) -> (2).

Training

The code repository includes data preparation and LoRA training for both tasks. Generate task-specific metadata from the training split, then run:

export MODEL_DIR="$PWD/models/Qwen-Image-Edit-2509"
export DATASET_ROOT=/path/to/Garments2Look-data
export TASK=inpainting
export METADATA="$PWD/data/metadata/train-inpainting.json"
export OUTPUT_DIR="$PWD/models/train/$TASK"
NPROC=1 SEED=123 bash scripts/train/train_lora.sh

# Editing: use editing-task metadata and a separate output directory.
export TASK=editing
export METADATA="$PWD/data/metadata/train-editing.json"
export OUTPUT_DIR="$PWD/models/train/$TASK"
NPROC=1 SEED=123 bash scripts/train/train_lora.sh

Defaults: rank 32, learning rate 1e-4, two epochs, gradient checkpointing, and a 1,048,576-pixel budget. For data preparation, multi-GPU usage, and metadata fields, see the code README. Editing-task source images are available in the updated dataset. The bundled examples are test samples and should not be used as a training benchmark.

Limitations and license

Fine accessory details, garment fidelity, styling, and pose preservation may vary. The example is illustrative, not an aggregate evaluation. The base model is required, and its license applies separately. An explicit adapter license has not yet been specified in this model repository.

Citation

If you use our dataset or models in your research, please consider citing our paper:

@inproceedings{cvpr2026garments2look,
    title={Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories},
    author={Hu, Junyao and Cheng, Zhongwei and Wong, Waikeung and Zou, Xingxing},
    booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    year={2026}
}
Downloads last month
-
Inference Providers NEW

Model tree for ArtmeScienceLab/Garments2Look-LoRA

Adapter
(86)
this model

Paper for ArtmeScienceLab/Garments2Look-LoRA