Instructions to use whosouravsharma/diffusiondb-sd15-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use whosouravsharma/diffusiondb-sd15-lora with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("stable-diffusion-v1-5/stable-diffusion-v1-5", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("whosouravsharma/diffusiondb-sd15-lora") prompt = "a anthropomorphic lion wizard, diffuse lighting, fantasy, intricate, elegant, highly detailed, lifelike, photorealistic, digital painting, artstation, illustration, concept art, smooth, sharp focus, naturalism, trending on byron's - muse, by greg rutkowski and greg staples" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
|
Download README.md from whosouravsharma/diffusiondb-sd15-lora: direct link, hf CLI and curl.
- Browser
- Download file 19.1 kB
-
https://huggingface.co/whosouravsharma/diffusiondb-sd15-lora/resolve/main/README.md
- Command line
-
hf download hf://whosouravsharma/diffusiondb-sd15-lora/README.md
-
curl -L -o README.md https://huggingface.co/whosouravsharma/diffusiondb-sd15-lora/resolve/main/README.md
19.1 kB
| base_model: stable-diffusion-v1-5/stable-diffusion-v1-5 | |
| base_model_relation: adapter | |
| library_name: diffusers | |
| pipeline_tag: text-to-image | |
| license: creativeml-openrail-m | |
| datasets: | |
| - whosouravsharma/text-to-image-diffusiondb-2M | |
| tags: | |
| - text-to-image | |
| - stable-diffusion | |
| - stable-diffusion-diffusers | |
| - lora | |
| - diffusers | |
| thumbnail: https://huggingface.co/whosouravsharma/diffusiondb-sd15-lora/resolve/main/images/pair-002.jpg | |
| widget: | |
| - text: a anthropomorphic lion wizard, diffuse lighting, fantasy, intricate, elegant, highly detailed, lifelike, photorealistic, digital painting, artstation, illustration, concept art, smooth, sharp focus, naturalism, trending on byron's - muse, by greg rutkowski and greg staples | |
| output: | |
| url: images/pair-002.jpg | |
| - text: a painting of a happy frog under the rain wearing a rainy coat by kazuo oga | |
| output: | |
| url: images/pair-006.jpg | |
| - text: mushrooms growing through a human skull in the woods | |
| output: | |
| url: images/pair-036.jpg | |
| - text: robots creating new robots in a factory, oil painting by justin gerard, deviantart, hd, 8 k | |
| output: | |
| url: images/pair-015.jpg | |
| - text: a painting of a silhouette in the water in front of a eclipse!!! with a red sky, by jeffrey smith, noah bradley, peter mohrbacher, behance contest winner, symbolism, darksynth, poster art, apocalypse art, hellish background | |
| output: | |
| url: images/pair-033.jpg | |
| - text: a giant skull with intricate rune carvings and glowing eyes with symmetrically braided lovecraftian tentacles haunting the cosmos by dan mumford, twirling smoke trail, a twisting vortex of dying galaxies, digital art, vivid colors, highly detailed | |
| output: | |
| url: images/pair-049.jpg | |
| model-index: | |
| - name: diffusiondb-sd15-lora (checkpoint-4240) | |
| results: | |
| - task: | |
| type: text-to-image | |
| dataset: | |
| name: DiffusionDB held-out prompts (report split, 506 images) | |
| type: whosouravsharma/text-to-image-diffusiondb-2M | |
| split: validation | |
| revision: v2-clean | |
| metrics: | |
| - name: KID (×10³, vs. real held-out images) | |
| type: kid | |
| value: 2.88 | |
| - name: FID (vs. real held-out images) | |
| type: fid | |
| value: 98.0 | |
| - name: CLIP score (ViT-L/14) | |
| type: clip_score | |
| value: 28.98 | |
| # diffusiondb-sd15-lora | |
| A rank-32 **style LoRA** for Stable Diffusion 1.5, trained on 13,598 prompt–image | |
| pairs from [DiffusionDB](https://huggingface.co/datasets/poloclub/diffusiondb). | |
| It gives SD 1.5 a **more saturated, higher-contrast, illustrative look** while | |
| following prompts just as well as the base model. | |
| <Gallery /> | |
| *Each image: plain SD 1.5 on the left, the same prompt and seed with the LoRA on | |
| the right.* | |
| | | | | |
| |---|---| | |
| | **What it does** | Adds bolder colour, stronger contrast and more dramatic lighting. Painterly prompts come out more like illustration | | |
| | **Prompt-following** | Unchanged: CLIP score 28.98 vs. 28.94 for base | | |
| | **Base model** | [`stable-diffusion-v1-5/stable-diffusion-v1-5`](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5) | | |
| | **Recommended checkpoint** | `checkpoints/checkpoint-4240` (25.5 MB) | | |
| | **Trigger word** | None. It applies to every prompt | | |
| | **Try it** | [DiffusionDB SD 1.5 LoRA Space](https://huggingface.co/spaces/whosouravsharma/diffusiondb-sd15-lora). The strength slider gives a live before/after | | |
| | **License** | CreativeML OpenRAIL-M | | |
| > **Finding worth knowing.** The adapter was trained on DiffusionDB in order to | |
| > reproduce DiffusionDB's look. It doesn't. DiffusionDB is itself Stable | |
| > Diffusion 1.x output, and plain SD 1.5 already matches it statistically | |
| > (KID ≈ 0). The LoRA instead adds a distinct style of its own, and measurably | |
| > moves outputs *away* from DiffusionDB. See [Evaluation](#evaluation). | |
| ## Table of contents | |
| - [Model details](#model-details) | |
| - [Uses](#uses) | |
| - [How to get started](#how-to-get-started) | |
| - [Evaluation](#evaluation) | |
| - [Bias, risks, and limitations](#bias-risks-and-limitations) | |
| - [Training details](#training-details) | |
| - [Environmental impact](#environmental-impact) | |
| - [Technical specifications](#technical-specifications) | |
| - [Reproduce](#reproduce) | |
| - [Citation](#citation) | |
| ## Model details | |
| | | | | |
| |---|---| | |
| | **Developed by** | [whosouravsharma](https://huggingface.co/whosouravsharma) | | |
| | **Model type** | LoRA adapter for a latent diffusion text-to-image model | | |
| | **Base model** | `stable-diffusion-v1-5/stable-diffusion-v1-5` (trained under its former name, `runwayml/stable-diffusion-v1-5`) | | |
| | **Adapter** | LoRA, rank 32, alpha 32, on the UNet attention projections (`to_q`, `to_k`, `to_v`, `to_out.0`) | | |
| | **Trainable parameters** | About 6.4M, against a UNet of about 860M | | |
| | **Language** | English prompts | | |
| | **License** | CreativeML OpenRAIL-M, the same as the base model | | |
| | Resource | Link | | |
| |---|---| | |
| | Demo | [Space: diffusiondb-sd15-lora](https://huggingface.co/spaces/whosouravsharma/diffusiondb-sd15-lora) | | |
| | Inference backend | [Space: diffusiondb-sd15-lora-inference](https://huggingface.co/spaces/whosouravsharma/diffusiondb-sd15-lora-inference) | | |
| | Training data | [whosouravsharma/text-to-image-diffusiondb-2M](https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M) | | |
| | Training and evaluation code | [`training/`](./training) and [`eval/`](./eval) in this repository | | |
| ## Uses | |
| ### Direct use | |
| - Giving SD 1.5 images a punchier, more saturated, illustrative finish, at an | |
| adjustable strength. | |
| - Studying how a fine-tune differs from its base. At strength 0 the output is | |
| pixel-identical to plain SD 1.5 at the same seed, so every difference comes | |
| from the adapter. | |
| - As a documented example of a LoRA fine-tune evaluated end to end: a | |
| leak-free split, a fixed-seed evaluation set, distribution metrics, a safety | |
| check and a memorization check. | |
| ### Out-of-scope use | |
| - Any use that the CreativeML OpenRAIL-M license prohibits. | |
| - Reproducing the DiffusionDB distribution. Plain SD 1.5 already does that | |
| better (see [Evaluation](#evaluation)). | |
| - Muted, pastel or low-contrast palettes. The adapter tends to override them | |
| (see [Failure cases](#failure-cases)). | |
| - Resolutions other than 512×512. | |
| - Photorealistic images of real, identifiable people. | |
| - Any setting where unfiltered output reaches users with no moderation step. | |
| ## How to get started | |
| ```python | |
| import torch | |
| from diffusers import StableDiffusionPipeline | |
| pipe = StableDiffusionPipeline.from_pretrained( | |
| "stable-diffusion-v1-5/stable-diffusion-v1-5", | |
| torch_dtype=torch.float16, | |
| variant="fp16", | |
| ).to("cuda") | |
| pipe.load_lora_weights( | |
| "whosouravsharma/diffusiondb-sd15-lora", | |
| subfolder="checkpoints/checkpoint-4240", | |
| weight_name="pytorch_lora_weights.safetensors", | |
| adapter_name="diffusiondb", | |
| ) | |
| pipe.set_adapters(["diffusiondb"], adapter_weights=[1.0]) # 0.0 = plain SD 1.5 | |
| image = pipe( | |
| "a steampunk owl inside a glass jar, intricate detail", | |
| num_inference_steps=25, | |
| guidance_scale=7.5, | |
| generator=torch.Generator("cuda").manual_seed(42), | |
| ).images[0] | |
| image.save("owl.png") | |
| ``` | |
| This snippet is run verbatim as part of the evaluation (diffusers 0.31.0, | |
| torch 2.4.0). | |
| **Strength.** 1.0 gives the full effect. Around 0.5 keeps the style while toning | |
| down the colour when it gets too strong. 1.5 exaggerates the style and starts to change the | |
| composition. | |
|  | |
| *The same seed at strength 0, 0.5, 1.0 and 1.5. Strength 0 is plain SD 1.5.* | |
| ## Evaluation | |
| Everything below comes from one evaluation run on a single A10G (1 h 53 min). | |
| The full outputs, including every render and metric, can be reproduced with | |
| [`eval/eval_job.py`](./eval/eval_job.py). | |
| ### Protocol | |
| - **Held-out data.** The 1,000 validation images were split by prompt group, | |
| so the model never trained on these prompts. Those 1,000 were then split | |
| again by prompt hash: a **selection half** (494 images) to choose a | |
| checkpoint, and a **report half** (506 images) for the numbers below. The | |
| reported numbers are therefore free of selection bias. | |
| - **Rendering.** For every model: 30 steps, guidance 7.5, the pipeline's | |
| default PNDM scheduler, and seed 42 + the validation row index. Every model | |
| sees the same prompts and seeds. | |
| - **Metrics.** | |
| - **KID and FID** against the real held-out images (centre-cropped to 512, | |
| the same view as training). Lower means closer to DiffusionDB. KID is the | |
| primary metric because FID is biased at this sample size. | |
| - **CLIP score** (ViT-L/14) of each image against its own prompt, which | |
| measures prompt-following. | |
| ### Results | |
| | Model (report half, n = 506) | KID ×10³ ↓ | FID ↓ | CLIP score | | |
| |---|---|---|---| | |
| | SD 1.5 base | **0.14 ± 0.32** | **95.4** | 28.94 | | |
| | + LoRA `checkpoint-2000` | 2.87 ± 0.64 | 98.7 | 28.79 | | |
| | + LoRA `checkpoint-4240` | 2.88 ± 0.61 | 98.0 | 28.98 | | |
| | Real held-out images (reference) | — | — | 28.95 | | |
| *KID is the mean ± standard deviation over 100 subsets of 253 images each. | |
| CLIP scores have a standard deviation of about 3.9 per image, which makes the | |
| standard error of each mean about 0.17.* | |
| **What the numbers say** | |
| 1. **Base SD 1.5 already matches DiffusionDB.** Its KID is indistinguishable | |
| from zero. That's expected, because DiffusionDB was generated with SD 1.x. | |
| 2. **The LoRA moves outputs away from DiffusionDB.** KID rises by about | |
| 2.7×10⁻³, roughly four times the combined spread of the two KID estimates. This is the style shift visible in | |
| the gallery, and it is not an approach to the training distribution. | |
| 3. **Prompt-following is unchanged.** All CLIP scores are within noise of each | |
| other and of the real images. | |
| 4. **Checkpoints from step 2,000 onward are statistically identical.** On the | |
| selection half, step 2,000 won narrowly (KID 2.50 vs. 2.62 for step 4,240 | |
| and 2.90 for step 3,000, all ± about 0.6). On the report half the two are | |
| tied. `checkpoint-4240` stays the recommended checkpoint. | |
| **A likely cause, not yet tested.** The added saturation looks like | |
| classifier-free guidance overshooting. The adapter may have sharpened the gap | |
| between the conditional and unconditional predictions, so guidance 7.5 | |
| behaves like a higher value. Sweeping the guidance scale with the adapter would | |
| confirm or rule this out. | |
| ### Validation loss | |
|  | |
| This is noise-prediction MSE on the 1,000 held-out images, with identical noise | |
| and timesteps for every model. The loss falls 0.9% below the base model | |
| (0.1496 → 0.1482), mostly in the first 2,000 steps, and is flat after about | |
| 3,000. There's no sign of divergence or overfitting. As is common for diffusion | |
| fine-tunes, a lower denoising loss did not translate into samples closer to the | |
| data (see KID above). | |
| ### How the style develops during training | |
|  | |
| *The same prompts and seeds at every checkpoint. Most of the style is in place | |
| by step 1,500. Later checkpoints refine it rather than change it.* | |
| ### Failure cases | |
|  | |
| *"shattering of the moon's surface, digital art, illustration". The detailed | |
| scene collapses into a flat, sparse composition.* | |
|  | |
| *"glass vodka bottle by shusei nagaoka, kaws, david rudnick, airbrush on | |
| canvas, **pastel colors**, cell-shaded, 8 k". The adapter's saturation | |
| overrides the requested pastel palette.* | |
|  | |
| *"long distance shot of a tiny cute polar bear on a tiny iceberg in the middle | |
| of the ocean, sunset, atmospheric, hazy". The animal's anatomy degrades, and | |
| the hazy mood is lost.* | |
| ### Correctness checks | |
| | Check | Result | | |
| |---|---| | |
| | Strength 0 equals plain SD 1.5 (same seed) | ✅ Max pixel difference 0 | | |
| | Unloading the adapter restores plain SD 1.5 | ✅ Max pixel difference 0 | | |
| | Same seed gives the same image | ✅ Max pixel difference 0 | | |
| | The adapter changes the output | ✅ Mean pixel difference 41.7 / 255 | | |
| | The usage snippet above runs as written | ✅ | | |
| ## Bias, risks, and limitations | |
| ### Safety | |
| The SD safety checker was run afterwards over the report-half renders; it is | |
| disabled during generation in the demo. | |
| | Images (n = 506) | Flagged | | |
| |---|---| | |
| | Real held-out images | 7 (1.4%) | | |
| | SD 1.5 base | 13 (2.6%) | | |
| | + LoRA `checkpoint-2000` | 10 (2.0%) | | |
| | + LoRA `checkpoint-4240` | 13 (2.6%) | | |
| The adapter doesn't raise the flag rate over the base model. The counts are | |
| small, so treat these as rough rates. The training data was filtered with | |
| DiffusionDB's own NSFW scores (below 0.2), but that classifier is noisy, so the | |
| filtering is not a guarantee. | |
| ### Memorization | |
| For each report-half render, we found the most similar of the 13,598 training | |
| images by CLIP ViT-L/14 image-embedding cosine. Real held-out images, which | |
| have different prompts but the same style, give the baseline for how similar | |
| two unrelated DiffusionDB images can be. | |
| | Images (n = 506) | Median nearest-neighbour cosine | ≥ 0.95 | Max | | |
| |---|---|---|---| | |
| | Real held-out images | 0.845 | 12 | 0.995 | | |
| | SD 1.5 base | 0.833 | 3 | 0.971 | | |
| | + LoRA `checkpoint-2000` | 0.829 | 2 | 0.959 | | |
| | + LoRA `checkpoint-4240` | 0.833 | 2 | 0.961 | | |
| The LoRA's renders are no closer to the training images than the base model's. | |
|  | |
| *The 8 closest pairs for `checkpoint-2000`. They share a subject and style, not | |
| a composition, so none is a copy. CLIP similarity is a proxy for copying, not | |
| proof either way.* | |
| ### Limitations | |
| - **The style is fixed.** Saturation and contrast go up on every prompt, | |
| including ones that ask for muted or pastel colours. | |
| - **Occasional composition loss.** Some detailed scenes are simplified (see | |
| [Failure cases](#failure-cases)). | |
| - **Small training set.** 13,598 images is enough to fine-tune, not to train | |
| from scratch. | |
| - **It inherits SD 1.x flaws.** The training images were themselves SD 1.x | |
| outputs, including their artifacts. | |
| - **Centre crop only.** There was no aspect-ratio bucketing, and roughly half | |
| the source images aren't square, so their edges were lost in training. | |
| - **The text encoder was frozen,** so the model understands prompts exactly as | |
| SD 1.5 does. | |
| ### Risks and recommendations | |
| - **Turn the safety checker back on.** The demo Space disables it. Put a | |
| safety checker or other moderation step in front of any public-facing use. | |
| - **It inherits SD 1.5's biases.** SD 1.5 and its LAION training data carry | |
| social and cultural biases. This adapter doesn't reduce them. | |
| - **Artist names in prompts.** Many DiffusionDB prompts name living artists, | |
| and the adapter learned from those prompt–image pairs. | |
| ## Training details | |
| ### Training data | |
| [`whosouravsharma/text-to-image-diffusiondb-2M`](https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M) | |
| at revision `v2-clean`: 13,598 train and 1,000 validation images, built from | |
| parts 1–20 of [`poloclub/diffusiondb`](https://huggingface.co/datasets/poloclub/diffusiondb). | |
| The validation split is made **by prompt group**: DiffusionDB users often ran | |
| the same prompt at many seeds, so every image sharing a normalized prompt goes | |
| to the same side. | |
| <details> | |
| <summary>Filtering, stage by stage (from the dataset's <code>manifest.json</code>)</summary> | |
| | Stage | Rows kept | Removed | | |
| |---|---|---| | |
| | Source metadata | 20,000 | — | | |
| | Prompt has at least 4 words | 19,513 | 487 | | |
| | Short side ≥ 384 px and area ≥ 262,144 px | 19,426 | 87 | | |
| | Image and prompt NSFW scores below 0.2 | 15,451 | 3,975 | | |
| | At most 2 images per normalized prompt | 14,606 | 845 | | |
| | Exact SHA-256 duplicates removed | **14,598** | 8 | | |
| </details> | |
| ### Training procedure | |
| Every image was encoded once with | |
| [`stabilityai/sd-vae-ft-mse`](https://huggingface.co/stabilityai/sd-vae-ft-mse) | |
| after resizing and a centre crop to 512×512. The results are stored as fp16 | |
| latents: both the posterior mean and its log-variance, so each training step | |
| samples a fresh latent. The text embeddings were not cached, which keeps 10% | |
| caption dropout possible for classifier-free guidance. | |
| <details> | |
| <summary>Hyperparameters</summary> | |
| | | | | |
| |---|---| | |
| | Objective | ε-prediction (MSE on the noise), DDPM noise schedule | | |
| | Epochs | 10 (424 optimizer steps per epoch, 4,240 in total) | | |
| | Batch size | 8 per step × 4 gradient-accumulation steps = 32 | | |
| | Optimizer | AdamW, lr 1e-4, weight decay 1e-2, grad-norm clip 1.0 | | |
| | Schedule | 500 linear warmup steps, then cosine decay | | |
| | Caption dropout | 10% | | |
| | LoRA init | Gaussian | | |
| | Precision | Frozen weights in fp16; LoRA weights in fp32; forward pass under fp16 autocast | | |
| | Seed | 42 | | |
| | Checkpoints | Every 500 steps, plus the final step (4,240) | | |
| Each checkpoint has a `state.json` recording the exact values it was trained | |
| with. | |
| </details> | |
| ## Environmental impact | |
| | | | | |
| |---|---| | |
| | Hardware | 1× NVIDIA A10G (Hugging Face Jobs, `a10g-small`) | | |
| | Training | 2.68 h (plus 0.25 h caching latents and 0.04 h baseline samples) | | |
| | Evaluation | 1.88 h | | |
| | Total | About 4.9 GPU-hours | | |
| | Carbon emitted | Not measured | | |
| ## Technical specifications | |
| SD 1.5 latent diffusion: CLIP ViT-L/14 text encoder (frozen), a UNet of about | |
| 860M parameters (frozen, with LoRA on the attention projections), and a VAE | |
| with 8× downsampling. Every stage runs as a standalone PEP 723 script on | |
| Hugging Face Jobs. Each job pulls its inputs from the Hub, and training can | |
| resume from any checkpoint with `RESUME_FROM`. | |
| <details> | |
| <summary>Repository layout</summary> | |
| ``` | |
| checkpoints/ | |
| checkpoint-<step>/ step = 500, 1000, … 4000, 4240 | |
| pytorch_lora_weights.safetensors the adapter | |
| optimizer.pt optimizer and grad-scaler state, for resuming | |
| state.json step, epoch, hyperparameters | |
| samples/base/ plain SD 1.5 on the 50 eval prompts | |
| images/ figures used in this card | |
| training/ data caching, training and sampling scripts | |
| eval/eval_job.py the evaluation that produced the numbers above | |
| ``` | |
| </details> | |
| ## Reproduce | |
| ```bash | |
| # training (run from training/) | |
| python3 main.py baseline # plain SD 1.5 samples, for comparison | |
| python3 main.py latents # VAE-encode the dataset once | |
| python3 main.py train # LoRA fine-tune | |
| python3 main.py sample checkpoint-4240 | |
| # evaluation: needs no token, writes only to the mounted bucket | |
| hf jobs uv run --flavor a10g-small --timeout 4h \ | |
| -v hf://buckets/<you>/<bucket>:/out eval/eval_job.py | |
| ``` | |
| ## Citation | |
| Training data comes from DiffusionDB (CC0 1.0): | |
| ```bibtex | |
| @article{wangDiffusionDBLargescalePrompt2022, | |
| title = {DiffusionDB: A Large-Scale Prompt Gallery Dataset for Text-to-Image Generative Models}, | |
| author = {Wang, Zijie J. and Montoya, Evan and Munechika, David and Yang, Haoyang and Hoover, Benjamin and Chau, Duen Horng}, | |
| journal = {arXiv:2210.14896 [cs]}, | |
| year = {2022}, | |
| url = {https://arxiv.org/abs/2210.14896} | |
| } | |
| ``` | |
| ## Model card contact | |
| Open a discussion in the | |
| [Community tab](https://huggingface.co/whosouravsharma/diffusiondb-sd15-lora/discussions). | |