Instructions to use metazlb/MT-OPSD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use metazlb/MT-OPSD with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("fill-in-base-model", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("metazlb/MT-OPSD") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
MT-OPSD checkpoints
LoRA checkpoints for MT-OPSD: On-Policy Self-Distillation for Multi-Turn Image Editing. MT-OPSD trains an editor on its own self-generated multi-turn states, which keeps it following instructions and keeps its images intact over long editing sessions.
| Subfolder | Base model | LME-Bench SR@10 (base β MT-OPSD) | CR@10 (base β MT-OPSD) |
|---|---|---|---|
qwen-image-edit-2511/ |
Qwen/Qwen-Image-Edit-2511 | 0.03 β 0.44 | 0.55 β 0.02 |
firered-image-edit-1.0/ |
FireRedTeam/FireRed-Image-Edit-1.0 | 0.15 β 0.52 | 0.61 β 0.03 |
flux2-klein-base-9b/ |
black-forest-labs/FLUX.2-klein-base-9B | 0.12 β 0.38 | 0.25 β 0.04 |
Each subfolder holds pytorch_lora_weights.safetensors (rank 32, alpha 64, fp32, diffusers
format) and adapter_config.json (target modules and the sampling settings used in
training). Use of each LoRA is also subject to the license of its base model.
Usage
git clone https://github.com/liangbingzhao/MT-OPSD.git && cd MT-OPSD
pip install -r requirements.txt
hf download metazlb/MT-OPSD --include "qwen-image-edit-2511/*" --local-dir checkpoints
python inference.py --ckpt checkpoints/qwen-image-edit-2511 --image input.png --output_dir outputs/demo \
--instructions "Make it a snowy winter scene." "Add a red scarf to the dog." "Convert to a pencil sketch."
Use the settings the checkpoints were trained with (inference.py reads them from
adapter_config.json): a 512Γ512 working area (aspect preserved), 30 sampling steps and true
CFG 4.0. Keep the LoRA in fp32 with bf16 autocast, as inference.py does; casting the LoRA
to bf16 measurably weakens long-horizon robustness.
Citation
@article{zhao2026mtopsd,
title={MT-OPSD: On-Policy Self-Distillation for Multi-Turn Image Editing},
author={Zhao, Liangbing and Zhuo, Le and Elhoseiny, Mohamed},
year={2026}
}
- Downloads last month
- -