MEND SD3.5-M Three Reward

Official evaluation adapter from MEND: RL for Flow Models via Proximal Velocity Matching, trained on PickScore + HPSv2.1 + CLIPScore for 300 updates.

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik.

Paper · Code · Project page · Blog · Collection · HF paper

MEND selects proposed reward-gradient moves using a capped reward minus a quadratic displacement price, then fits the resulting velocity targets. The unchanged sample is always a candidate. Training uses an EMA behavior policy, with no KL term or frozen reference model in the loss.

Same-prompt, same-noise comparison: SD3.5-M above and MEND PickScore at 100 updates below

The comparison above is the paper's PickScore-100 example for both model cards; it does not show the three-reward checkpoint.

Download and generate

Accept the SD3.5 Medium base model terms and authenticate with hf auth login. The adapter is publicly downloadable; the base pipeline is downloaded separately. Install a CUDA PyTorch build suited to your system and the MEND package following installation instructions.

Use the paper's generation pipeline:

git clone https://github.com/shreshthsaini/MEND-RL.git
cd MEND-RL
uv pip install -e .
python scripts/generate.py --lora mend_open3 \
  --prompt "a small blue book on a large red book" --seeds 0 \
  --guidance_scale 1.0 --num_steps 40 --resolution 512 --batch_size 1 \
  --out_dir outputs/mend_open3

The official alias pins the original PEFT weights to a Hub commit and verifies their SHA256 checksums. Download without loading a model using python scripts/download_weights.py mend_open3. See the inference guide for offline use and revision pinning.

Or use Diffusers directly with the included pytorch_lora_weights.safetensors:

import torch
from diffusers import StableDiffusion3Pipeline

pipe = StableDiffusion3Pipeline.from_pretrained(
    "stabilityai/stable-diffusion-3.5-medium", torch_dtype=torch.bfloat16,
)
pipe.load_lora_weights("shreshthsaini/MEND-SD3.5M-ThreeReward", weight_name="pytorch_lora_weights.safetensors")
pipe.enable_model_cpu_offload()
image = pipe(
    "a small blue book on a large red book",
    height=512, width=512, num_inference_steps=40, guidance_scale=1.0,
    generator=torch.Generator(device="cpu").manual_seed(0),
).images[0]
image.save("mend.png")

Diffusers and the repository's evaluation sampler need not produce pixel-identical images. Use the repository pipeline to reproduce its evaluation settings. GPU memory depends on precision and offload settings; no minimum VRAM is claimed.

Checkpoint and evaluation

Setting Value
Base SD3.5 Medium
Training rewards PickScore + HPSv2.1 + CLIPScore
Updates 300
Adapter Evaluation EMA policy, rank 32, original alpha 64
Sampling 512 × 512, 40 deterministic Euler steps, guidance 1.0
Format Original PEFT adapter and equivalent Diffusers LoRA

Paper Table 1 uses DrawBench, 200 prompts × 5 seeds. Values below are the paper's original measurements, rather than a new release evaluation.

PickScore HPSv2.1 HPSv3 ImageReward CLIPScore Aesthetic DreamSim distance to base
23.89 0.344 5.87 1.28 0.300 5.92 0.488

Distance compares with base images generated from the same prompt, noise and sampling settings. Each configuration has one training seed. Optimizing these rewards does not guarantee improvement for every prompt or evaluator. See the paper for protocols and limitations. Other released MEND adapter.

Files and integrity

adapter_config.json and adapter_model.safetensors are the original evaluation adapter. pytorch_lora_weights.safetensors converts it to Diffusers by absorbing the original alpha/rank scale into LoRA B. Load one format at a time. Optimizer state, training checkpoints, the auxiliary behavior adapter, and base weights are excluded. release_manifest.json describes the release; SHA256SUMS hashes its files.

Release validation checks original hashes, all tensor shapes against the base architecture, PEFT loading, and numerical equivalence of the converted adapter. It does not rerun full image generation or paper metrics.

License and citation

Powered by Stability AI. Weights are subject to the Stability AI Community License, including its use and distribution conditions. The MEND code is Apache-2.0. Base models, reward models and datasets retain their own terms.

If you use these weights or MEND, please cite:

@misc{saini2026mend,
  title         = {{MEND}: {RL} For Flow Models via Proximal Velocity Matching},
  author        = {Saini, Shreshth and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu and Bovik, Alan C.},
  year          = {2026},
  eprint        = {2610.05954},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  doi           = {10.48550/arXiv.2610.05954},
  url           = {https://arxiv.org/abs/2610.05954}
}
Downloads last month
414
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shreshthsaini/MEND-SD3.5M-ThreeReward

Adapter
(113)
this model

Space using shreshthsaini/MEND-SD3.5M-ThreeReward 1

Collection including shreshthsaini/MEND-SD3.5M-ThreeReward

Paper for shreshthsaini/MEND-SD3.5M-ThreeReward