Instructions to use shreshthsaini/MEND-SD3.5M-ThreeReward with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use shreshthsaini/MEND-SD3.5M-ThreeReward with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("stabilityai/stable-diffusion-3.5-medium", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("shreshthsaini/MEND-SD3.5M-ThreeReward") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - PEFT
How to use shreshthsaini/MEND-SD3.5M-ThreeReward with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
MEND SD3.5-M Three Reward
Official evaluation adapter from MEND: RL for Flow Models via Proximal Velocity Matching, trained on PickScore + HPSv2.1 + CLIPScore for 300 updates.
Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik.
Paper · Code · Project page · Blog · Collection · HF paper
MEND selects proposed reward-gradient moves using a capped reward minus a quadratic displacement price, then fits the resulting velocity targets. The unchanged sample is always a candidate. Training uses an EMA behavior policy, with no KL term or frozen reference model in the loss.
The comparison above is the paper's PickScore-100 example for both model cards; it does not show the three-reward checkpoint.
Download and generate
Accept the SD3.5 Medium base model terms and authenticate with hf auth login. The adapter is publicly downloadable; the base pipeline is downloaded separately. Install a CUDA PyTorch build suited to your system and the MEND package following installation instructions.
Use the paper's generation pipeline:
git clone https://github.com/shreshthsaini/MEND-RL.git
cd MEND-RL
uv pip install -e .
python scripts/generate.py --lora mend_open3 \
--prompt "a small blue book on a large red book" --seeds 0 \
--guidance_scale 1.0 --num_steps 40 --resolution 512 --batch_size 1 \
--out_dir outputs/mend_open3
The official alias pins the original PEFT weights to a Hub commit and verifies their SHA256 checksums. Download without loading a model using python scripts/download_weights.py mend_open3. See the inference guide for offline use and revision pinning.
Or use Diffusers directly with the included pytorch_lora_weights.safetensors:
import torch
from diffusers import StableDiffusion3Pipeline
pipe = StableDiffusion3Pipeline.from_pretrained(
"stabilityai/stable-diffusion-3.5-medium", torch_dtype=torch.bfloat16,
)
pipe.load_lora_weights("shreshthsaini/MEND-SD3.5M-ThreeReward", weight_name="pytorch_lora_weights.safetensors")
pipe.enable_model_cpu_offload()
image = pipe(
"a small blue book on a large red book",
height=512, width=512, num_inference_steps=40, guidance_scale=1.0,
generator=torch.Generator(device="cpu").manual_seed(0),
).images[0]
image.save("mend.png")
Diffusers and the repository's evaluation sampler need not produce pixel-identical images. Use the repository pipeline to reproduce its evaluation settings. GPU memory depends on precision and offload settings; no minimum VRAM is claimed.
Checkpoint and evaluation
| Setting | Value |
|---|---|
| Base | SD3.5 Medium |
| Training rewards | PickScore + HPSv2.1 + CLIPScore |
| Updates | 300 |
| Adapter | Evaluation EMA policy, rank 32, original alpha 64 |
| Sampling | 512 × 512, 40 deterministic Euler steps, guidance 1.0 |
| Format | Original PEFT adapter and equivalent Diffusers LoRA |
Paper Table 1 uses DrawBench, 200 prompts × 5 seeds. Values below are the paper's original measurements, rather than a new release evaluation.
| PickScore | HPSv2.1 | HPSv3 | ImageReward | CLIPScore | Aesthetic | DreamSim distance to base |
|---|---|---|---|---|---|---|
| 23.89 | 0.344 | 5.87 | 1.28 | 0.300 | 5.92 | 0.488 |
Distance compares with base images generated from the same prompt, noise and sampling settings. Each configuration has one training seed. Optimizing these rewards does not guarantee improvement for every prompt or evaluator. See the paper for protocols and limitations. Other released MEND adapter.
Files and integrity
adapter_config.json and adapter_model.safetensors are the original evaluation adapter. pytorch_lora_weights.safetensors converts it to Diffusers by absorbing the original alpha/rank scale into LoRA B. Load one format at a time. Optimizer state, training checkpoints, the auxiliary behavior adapter, and base weights are excluded. release_manifest.json describes the release; SHA256SUMS hashes its files.
Release validation checks original hashes, all tensor shapes against the base architecture, PEFT loading, and numerical equivalence of the converted adapter. It does not rerun full image generation or paper metrics.
License and citation
Powered by Stability AI. Weights are subject to the Stability AI Community License, including its use and distribution conditions. The MEND code is Apache-2.0. Base models, reward models and datasets retain their own terms.
If you use these weights or MEND, please cite:
@misc{saini2026mend,
title = {{MEND}: {RL} For Flow Models via Proximal Velocity Matching},
author = {Saini, Shreshth and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu and Bovik, Alan C.},
year = {2026},
eprint = {2610.05954},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
doi = {10.48550/arXiv.2610.05954},
url = {https://arxiv.org/abs/2610.05954}
}
- Downloads last month
- 414
Model tree for shreshthsaini/MEND-SD3.5M-ThreeReward
Base model
stabilityai/stable-diffusion-3.5-medium