File size: 4,487 Bytes
edf3fc4
 
0086d5e
 
 
 
 
 
 
 
 
 
 
 
 
edf3fc4
0086d5e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38f12a8
0086d5e
 
 
68d6cf0
 
 
0086d5e
38f12a8
 
 
0086d5e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
547c368
 
 
 
 
0086d5e
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
---
license: apache-2.0
library_name: diffusers
pipeline_tag: text-to-image
base_model:
- stabilityai/stable-diffusion-3.5-medium
- Tongyi-MAI/Z-Image-Turbo
tags:
- diffusion
- text-to-image
- image-generation
- reinforcement-learning
- self-distillation
- lora
- arxiv:2608.24646
---

<div align="center">

# DiffusionOPSD: On-Policy Self-Distillation in Diffusion Models

**Reward-guided diffusion post-training through explicit, continually refreshed intermediate targets**

[![Paper](https://img.shields.io/badge/arXiv-2608.24646-b31b1b?logo=arxiv)](https://arxiv.org/abs/2608.24646)
[![Project Page](https://img.shields.io/badge/Project-Page-3B82F6)](https://diffusionopsd.github.io/)
[![Code](https://img.shields.io/badge/Code-GitHub-181717?logo=github)](https://github.com/worldbench/DiffusionOPSD)

<img src="assets/qualitative_gallery.jpg" width="100%" alt="Images generated with DiffusionOPSD">

</div>

## Overview

**DiffusionOPSD** is an on-policy self-distillation framework for reward-guided diffusion post-training. A frozen behavior policy collects on-policy denoising states and clean-output anchors; differentiable reward gradients construct bounded positive and negative targets around each anchor; and the trainable policy fits these detached targets before an EMA update refreshes the behavior policy.

By turning image-level rewards into explicit, continually refreshed intermediate supervision, DiffusionOPSD makes **target construction** and **finite realization** separately observable. Across SD3.5-M and Z-Image-Turbo, it achieves the best final held-out score in **19 of 20** reward-matched settings and reduces training GPU-hours relative to DiffusionNFT by **40%** and **63%**, respectively.

<p align="center">
  <img src="assets/method_overview.png" width="100%" alt="DiffusionOPSD method overview">
</p>

## Released Checkpoints

This repository provides nine rank-32 LoRA adapters:

| Checkpoint | Backbone | Training objective |
|---|---|---|
| [`sd35-m-hpsv2`](./sd35-m-hpsv2) | [Stable Diffusion 3.5 Medium](https://huggingface.co/stabilityai/stable-diffusion-3.5-medium) | HPSv2.1 |
| [`sd35-m-pickscore`](./sd35-m-pickscore) | [Stable Diffusion 3.5 Medium](https://huggingface.co/stabilityai/stable-diffusion-3.5-medium) | PickScore |
| [`sd35-m-clipscore`](./sd35-m-clipscore) | [Stable Diffusion 3.5 Medium](https://huggingface.co/stabilityai/stable-diffusion-3.5-medium) | CLIPScore |
| [`sd35-m-hpsv3`](./sd35-m-hpsv3) | [Stable Diffusion 3.5 Medium](https://huggingface.co/stabilityai/stable-diffusion-3.5-medium) | HPSv3 |
| [`z-image-turbo-hpsv2`](./z-image-turbo-hpsv2) | [Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | HPSv2.1 |
| [`z-image-turbo-pickscore`](./z-image-turbo-pickscore) | [Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | PickScore |
| [`z-image-turbo-clipscore`](./z-image-turbo-clipscore) | [Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | CLIPScore |
| [`z-image-turbo-hpsv3`](./z-image-turbo-hpsv3) | [Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | HPSv3 |
| [`z-image-turbo-pointwise`](./z-image-turbo-pointwise) | [Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | Pointwise reward |

Download all released adapters with:

```bash
hf download WeiChow/DiffusionOPSD --local-dir checkpoints/diffusionopsd
```

<p align="center">
  <img src="assets/training_curves.png" width="100%" alt="DiffusionOPSD training and held-out quality curves">
</p>

## Resources

- **Paper:** [On-Policy Self-Distillation in Diffusion Models](https://arxiv.org/abs/2608.24646)
- **Code:** [worldbench/DiffusionOPSD](https://github.com/worldbench/DiffusionOPSD)
- **Project page:** [diffusionopsd.github.io](https://diffusionopsd.github.io/)

Please refer to the [GitHub repository](https://github.com/worldbench/DiffusionOPSD) for installation, inference, evaluation, and training instructions.

## Citation

```bibtex
@article{zhou2026policy,
  title={On-Policy Self-Distillation in Diffusion Models},
  author={Zhou, Wei and Zhu, Xiongwei and Kong, Lingdong and Chen, Bo and Zhang, Lei and Liang, Yongyuan and Hou, Xiaoxia and Tian, Ye and Sun, Xian and Wang, Yingshuo and others},
  journal={arXiv preprint arXiv:2608.24646},
  year={2026}
}
```

## License

The released adapters are provided under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). Users must also comply with the licenses of the corresponding base models.