Instructions to use SykoSLM/SykoDiffusion-V1.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use SykoSLM/SykoDiffusion-V1.1 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("SykoSLM/SykoDiffusion-V1.1", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
File size: 13,329 Bytes
c470ded | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 | ---
library_name: diffusers
pipeline_tag: text-to-image
tags:
- diffusion
- text-to-image
- anime
- pixel-space
- unet
- classifier-free-guidance
datasets:
- huggan/anime-faces
---
# SykoDiffusion-V1.1
A small, text-conditioned **pixel-space diffusion model** that generates **64Γ64 anime faces** from short attribute prompts such as `anime face, purple long hair, purple eyes, smiling`.
The model was trained from scratch (~38M parameters) on a single Google TPU v5e-1. It has **no VAE**: the UNet works directly on RGB pixels and its output is the final image.
| | |
|---|---|
| **Task** | Text-to-image (structured attribute prompts) |
| **Output resolution** | 64Γ64 RGB |
| **Architecture** | `diffusers.UNet2DConditionModel`, ~38.4M parameters |
| **Diffusion space** | Pixel space (no VAE / no latent compression) |
| **Text encoder** | CLIP ViT-B/32 text encoder (frozen, `openai/clip-vit-base-patch32`) |
| **Sampler** | DDIM (Ξ· = 0) with classifier-free guidance |
---
## Table of contents
- [About the "no VAE" design](#about-the-no-vae-design)
- [Quick start](#quick-start)
- [Prompting guide](#prompting-guide)
- [What the model can and cannot do](#what-the-model-can-and-cannot-do)
- [Model details](#model-details)
- [Training details](#training-details)
- [Repository contents](#repository-contents)
- [Using the original script](#using-the-original-script)
- [Limitations, bias and responsible use](#limitations-bias-and-responsible-use)
- [License](#license)
- [Acknowledgements](#acknowledgements)
---
## About the "no VAE" design
Many popular text-to-image systems (for example Stable Diffusion) are *latent* diffusion models. They compress every image into a small latent tensor with a VAE, run the diffusion process on that latent, and use the VAE decoder at the end to turn the latent back into pixels.
**SykoDiffusion does not do this.**
```
Latent diffusion: noise β UNet β latent (compressed, not an image) β VAE decoder β image
SykoDiffusion: noise β UNet β RGB image (64Γ64Γ3)
```
- The UNet takes a 3-channel image as input and predicts 3-channel noise (`in_channels=3`, `out_channels=3`).
- Training images are scaled to `[-1, 1]` and fed straight into the UNet.
- There is nothing to decode. The sampler's final tensor *is* the picture.
At 64Γ64 the pixel-space approach is computationally affordable and avoids the blur and colour artefacts a VAE can introduce.
**Consequences worth knowing:**
- A VAE is not a super-resolution tool. Adding one to this model would not increase the output resolution.
- The output resolution is fixed at 64Γ64. To obtain larger images you would need a separate super-resolution model, which enlarges (and can amplify) any artefacts in the source image, or a model trained at a higher resolution.
---
## Quick start
The repository stores the UNet in the standard `diffusers` format, plus the CLIP tokenizer and text encoder. It is **not** a full `DiffusionPipeline`, so `DiffusionPipeline.from_pretrained(...)` will not work. The snippet below is a self-contained sampler that reproduces the original DDIM + classifier-free-guidance loop.
```bash
pip install torch diffusers transformers pillow numpy huggingface_hub
```
If the repository is private, authenticate first (`huggingface-cli login`, or pass `token=...` to each `from_pretrained` call).
```python
import numpy as np
import torch
from PIL import Image
from diffusers import DDIMScheduler, UNet2DConditionModel
from transformers import CLIPTextModelWithProjection, CLIPTokenizer
REPO = "SykoSLM/SykoDiffusion-V1.1"
RES, T_STEPS, TOK_LEN = 64, 1000, 24
device = "cuda" if torch.cuda.is_available() else "cpu"
unet = UNet2DConditionModel.from_pretrained(REPO).to(device).eval()
tok = CLIPTokenizer.from_pretrained(REPO, subfolder="tokenizer")
txt = CLIPTextModelWithProjection.from_pretrained(REPO, subfolder="text_encoder").to(device).eval()
@torch.no_grad()
def encode(prompts):
t = tok(prompts, padding="max_length", max_length=TOK_LEN, truncation=True, return_tensors="pt").to(device)
return txt(**t).last_hidden_state # (B, 24, 512)
@torch.no_grad()
def generate(prompt, n=6, steps=50, cfg=4.0, seed=0):
cond = encode([prompt]).repeat_interleave(n, dim=0)
null = encode([""]).expand(n, -1, -1) # empty prompt = unconditional branch
ctx = torch.cat([null, cond])
sched = DDIMScheduler(
num_train_timesteps=T_STEPS,
beta_schedule="squaredcos_cap_v2",
clip_sample=True,
prediction_type="epsilon",
timestep_spacing="trailing",
)
sched.set_timesteps(steps)
g = torch.Generator().manual_seed(seed)
x = torch.randn(n, 3, RES, RES, generator=g).to(device)
for t in sched.timesteps:
eps = unet(torch.cat([x, x]), t.to(device), encoder_hidden_states=ctx).sample
e_uncond, e_cond = eps.chunk(2)
eps = e_uncond + cfg * (e_cond - e_uncond) # classifier-free guidance
x = sched.step(eps, t, x).prev_sample
return (x.clamp(-1, 1) + 1) / 2 # (n, 3, 64, 64) in [0, 1]
imgs = generate("anime face, purple long hair, purple eyes, smiling", n=6, steps=50, cfg=4.0, seed=0)
imgs = (imgs.permute(0, 2, 3, 1).cpu().numpy() * 255 + 0.5).astype(np.uint8)
Image.fromarray(np.concatenate(list(imgs), axis=1)).save("samples.png")
```
**Sampling parameters**
| Parameter | Default | Notes |
|---|---|---|
| `steps` | 50 | DDIM steps. Going to 100 usually changes little. |
| `cfg` | 4.0 | Guidance scale. Higher values follow the prompt more strongly but can produce dark or oversaturated samples. It is worth experimenting with this value. |
| `seed` | 0 | Same seed + same settings gives the same images. |
---
## Prompting guide
The model was trained on **templated captions**, so it responds best to the same format:
```
anime face, <hair colour> <long|short> hair, <eye colour> eyes[, smiling]
```
**Vocabulary seen during training**
| Attribute | Values |
|---|---|
| Hair colour | blonde, brown, black, blue, pink, purple, green, red, silver, white, orange |
| Hair length | long, short |
| Eye colour | blue, red, green, brown, purple, yellow, orange, pink, black |
| Expression | `smiling` (appended only when a smile was detected) |
Examples:
```
anime face, blue long hair, red eyes, smiling
anime face, brown short hair, green eyes
anime face, pink long hair, purple eyes, smiling
```
Notes:
- Prompts are truncated to **24 CLIP tokens**.
- There is no "neutral" or "not smiling" tag. A neutral expression corresponds to leaving `smiling` out.
- Free-form prompts (styles, accessories, backgrounds, poses) were never seen in training and are not expected to work reliably.
- An empty prompt (`""`) is the unconditional branch used for classifier-free guidance.
---
## What the model can and cannot do
**What it can do**
- Generate 64Γ64 anime-style faces conditioned on hair colour, hair length, eye colour and (optionally) a smile.
- Produce varied faces for the same prompt by changing the seed.
- Run on CPU, GPU or TPU. The UNet is small enough for casual local experimentation.
**What it cannot do**
- **Produce anything larger than 64Γ64.** The resolution is fixed.
- **Understand free-form text.** Only the templated attribute vocabulary above is covered by training.
- **Generate anything other than anime faces.** No full-body characters, backgrounds, objects, text or other subjects.
- **Guarantee clean results.** Some samples show asymmetric or blurry eyes and mouths, or are noticeably dark and oversaturated. This happens in a fraction of samples and varies by seed and prompt.
- **Follow attributes perfectly.** The training labels were produced automatically (see below) and contain errors, so the conditioning is approximate.
- **Support negative prompts or image-to-image editing.** Neither was implemented or evaluated.
No quantitative evaluation (e.g. FID) has been run. Quality has only been assessed by visually inspecting fixed-seed sample grids during training.
---
## Model details
| | |
|---|---|
| Backbone | `UNet2DConditionModel` (from `diffusers`), trained from scratch |
| Parameters | ~38.4M |
| Input / output | 3 Γ 64 Γ 64 (RGB, range `[-1, 1]`) |
| Blocks | `DownBlock2D`, `CrossAttnDownBlock2D` Γ2 β mirrored up blocks (`CrossAttnUpBlock2D` Γ2, `UpBlock2D`) |
| Channels per level | 96, 192, 384 |
| Layers per block | 1 |
| Attention | Cross-attention to text at the 32Γ32 and 16Γ16 levels; no attention at 64Γ64 |
| Attention head dim | 8 |
| Group-norm groups | 32 |
| Text conditioning | CLIP ViT-B/32 text encoder, last hidden state, 24 tokens Γ 512 dims (frozen) |
| Prediction target | Ξ΅ (noise) |
| Noise schedule | `squaredcos_cap_v2`, 1000 training timesteps |
| Output layer | Zero-initialised at the start of training |
---
## Training details
**Data**
- Source: [`huggan/anime-faces`](https://huggingface.co/datasets/huggan/anime-faces).
- 43,102 images after removing unreadable files, centre-cropped to a square and resized to 64Γ64.
- **Captions are synthetic.** They were generated automatically with CLIP ViT-B/32 zero-shot classification (hair colour, hair length, eye colour, smile) and assembled into the templated format above. Because these labels are model predictions rather than human annotations, they contain noise.
**Setup**
| | |
|---|---|
| Hardware | Google TPU v5e-1 (Colab), PyTorch/XLA |
| Precision | bf16 autocast for compute; FP32 master weights, EMA and optimiser state |
| Steps | 40,000 (~59 epochs) |
| Batch size | 64 |
| Optimiser | AdamW (Ξ² = 0.9, 0.99; weight decay 0.01), gradient clipping at 1.0 |
| Learning rate | 2e-4 peak, 1,000-step linear warm-up, cosine decay to 5% of peak |
| Loss weighting | Min-SNR (Ξ³ = 5) |
| EMA | Decay 0.9995 (the released weights are the EMA weights) |
| Augmentation | Random horizontal flip |
| Conditioning dropout | 10% of captions replaced by the empty caption (enables classifier-free guidance) |
**A note on the loss.** The training loss plateaued around 0.025 for most of training even though sample quality kept improving. This is common for diffusion models: most of the loss is irreducible noise-prediction error, and the fine-detail improvements that matter visually contribute very little to it. Fixed-seed sample grids were a more reliable progress signal than the loss curve.
---
## Repository contents
```
.
βββ config.json # UNet configuration
βββ diffusion_pytorch_model.safetensors # EMA UNet weights
βββ text_encoder/ # CLIP ViT-B/32 text encoder (frozen)
βββ tokenizer/ # CLIP tokenizer
βββ anime_t2i_tpu.py # Full training / sampling script
βββ README.md
```
The CLIP vision encoder is not included. It was only used once, during data preparation, to produce the captions, and is not needed for generation.
---
## Using the original script
`anime_t2i_tpu.py` contains the complete pipeline: data preparation, training on TPU (PyTorch/XLA) and sampling. It reads the model from `<WORK_DIR>/ckpt/ema_unet`, so download the UNet files (`config.json` and `diffusion_pytorch_model.safetensors`) into that folder first.
```bash
# Sample from a trained model
python anime_t2i_tpu.py sample \
--prompt "anime face, blue long hair, red eyes, smiling" \
--n 6 --cfg 4 --steps 50 --seed 0
# Prepare data and train from scratch (requires a TPU runtime for the XLA path)
python anime_t2i_tpu.py prepare
python anime_t2i_tpu.py train --steps 40000 --batch 64 --lr 2e-4
```
If you retrain or resume, keep `block_out_channels` and `layers_per_block` in `build_unet` identical to the values listed under [Model details](#model-details), otherwise checkpoint loading will fail with a size mismatch.
---
## Limitations, bias and responsible use
- **Dataset bias.** The model reproduces the style and demographic distribution of its training dataset, including any imbalance in hair colours, eye colours, expressions and character appearance. Rare attribute combinations are generated less reliably.
- **Label noise.** Automatic CLIP labelling means the attribute conditioning is imperfect.
- **Low resolution.** 64Γ64 output is suited to experimentation, prototyping and research, not to production artwork.
- **Dataset licensing.** The training images come from a third-party dataset. Check the [dataset card](https://huggingface.co/datasets/huggan/anime-faces) for its terms before any commercial use of the model or its outputs.
- The model generates stylised illustrations of fictional faces. It is not intended to depict real people.
---
## License
[Specify the license for the model weights here.]
The CLIP text encoder and tokenizer included in this repository are from OpenAI's CLIP (`openai/clip-vit-base-patch32`), released under the MIT license.
---
## Acknowledgements
- [OpenAI CLIP](https://github.com/openai/CLIP): frozen text encoder and zero-shot captioning.
- [Hugging Face `diffusers`](https://github.com/huggingface/diffusers): UNet implementation and schedulers.
- [`huggan/anime-faces`](https://huggingface.co/datasets/huggan/anime-faces): training images.
- Techniques used: DDPM/DDIM sampling, classifier-free guidance, Min-SNR loss weighting, and EMA of weights.
|