Instructions to use oloruntobiolutola/facebench-students with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use oloruntobiolutola/facebench-students with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("oloruntobiolutola/facebench-students", dtype=torch.bfloat16, device_map="cuda") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
FaceBench students: one-step face editors distilled from Qwen-Image-Edit
Two small image-to-image models that perform eight binary face edits in a single forward pass. They were distilled from Qwen-Image-Edit (20B parameters, 50 steps, about 52 s per edit on an H100) as part of the FaceBench project, which benchmarks identity-preserving face editing.
SDXS student on a person not seen in training. From left to right: source, add glasses, remove glasses, add beard,
remove beard, smile, neutral, older, younger.
| Folder | Model | Parameters | Resolution | Speed |
|---|---|---|---|---|
sdxs/ |
SDXS student (recommended): SDXS-512-0.9 with LoRA merged in, gated one-step head, learned copy-back mask | 331M | 512×416 | 38 ms one at a time, 11 ms per image batched (L40S, fp16); about 0.3 s on an Apple M-series GPU |
pixel/ |
Pixel student: hybrid compact U-Net trained from scratch, predicts a residual on the source | 29M | 256×208 | 29 ms batched (L4) |
Edits
| Edit name | Meaning | Teacher prompt it was distilled from |
|---|---|---|
Eyeglasses+ |
add glasses | "Add eyeglasses." |
Eyeglasses- |
remove glasses | "Remove the eyeglasses." |
No_Beard+ |
add a beard | "Add a beard." |
No_Beard- |
remove the beard | "Remove the beard." |
Smiling+ |
smile | "Make the person smile." |
Smiling- |
neutral expression | "Change the facial expression to neutral, not smiling, with the lips closed." |
Young+ |
older | "Make the same person look older, keeping their face shape, gender, hairstyle and background unchanged apart from age." |
Young- |
younger | "Make the same person look like a young adult, keeping their face shape, skin tone, ethnicity, gender, hairstyle, accessories and background unchanged apart from age." |
The edits are fixed: the prompt embeddings are stored with the model (sdxs/editor.safetensors), so no text
encoder is needed. Free-text prompts are not supported.
Usage
pip install torch diffusers safetensors pillow huggingface_hub
hf download oloruntobiolutola/facebench-students --include "inference.py" --local-dir .
python inference.py --model sdxs --edit Eyeglasses+ --input face.jpg --out out/
python inference.py --model sdxs --edit all --input photos/ --out out/ --sheet # every edit, plus a sheet per photo
python inference.py --model pixel --edit Young+ --input face.jpg --out out/
inference.py downloads the weights it needs from this repo. From Python:
from PIL import Image
from inference import load
ed = load("sdxs", device="cuda") # or "pixel"; device defaults to cuda > mps > cpu
ed.edit(Image.open("face.jpg"), "Smiling+").save("smile.png")
Inputs should be face crops framed like CelebA aligned photos (face centred, about 178×218). Other photos are resized to the model's resolution, and results degrade the further they are from that framing.
How it works
SDXS student.
- The tiny VAE encodes the source.
- One U-Net step at t = 999 runs with the edit's fixed prompt embedding, which gives an x₀ estimate.
- A learned per-channel gate moves the source latent towards that estimate:
z = z_src + g·(x̂₀ − z_src). - The VAE decodes the result.
- A learned mask,
sigmoid(b − k·blur|D(z) − D(z_src)|), copies the original pixels back wherever the edit left the image unchanged. This keeps backgrounds sharp and removes ghosting on removals.
It was trained with LoRA (rank 32) on the U-Net, plus the VAE and conv_in, using L1 + LPIPS against the teacher's
edits.
Pixel student. A 29M U-Net with convolutional ResBlocks at high resolution, self-attention at 32×26 and a FiLM
edit embedding. It predicts output = source + f(source, edit), trained with L1 + LPIPS.
Training data. About 2,000 Qwen-Image-Edit edits of CelebA photos, with balanced edit types. Validation used people held out from training.
Evaluation
Closeness to the teacher. Measured at 256×208 on held-out people, against the teacher's own edit:
| Model | PSNR gain over copying the source | LPIPS ↓ |
|---|---|---|
| SDXS student | +3.66 dB | 0.051 |
| Pixel student | +2.94 dB | 0.061 |
Blind human evaluation. Two graders judged 40 requests (5 per edit type) on 40 people not in the training data. Each request was edited by six models, shown in random order with names hidden, at equal display size. Graders agreed reliably on whether the edit was done (Krippendorff's α = 0.81) and on same person (α = 0.77 ordinal, 0.82 interval).
| Model | Edit done | Done and same person ≥ 4/5 | Same person (1–5) | Quality (1–5) | Picked best |
|---|---|---|---|---|---|
| Qwen-Image-Edit (teacher) | 96% | 91% | 4.47 | 4.69 | 30% |
| SDXS student | 96% | 88% | 4.36 | 4.45 | 29% |
| FLUX.1 Kontext | 70% | 70% | 4.74 | 4.74 | 20% |
| LLaDA-Image | 86% | 51% | 3.27 | 4.20 | 6% |
| Pixel student | 56% | 52% | 3.44 | 3.09 | 10% |
| SDXS-512 untrained (zero-shot img2img) | 11% | 2% | 1.39 | 1.84 | 1% |
The SDXS student is not significantly different from its teacher on any measure (paired tests, Holm-corrected). It is ahead of FLUX.1 Kontext on success and of LLaDA-Image on identity.
The pixel student handles the local additions (glasses, beard, smile) but fails the age edits and glasses removal.
All trained checkpoints (checkpoints/)
Every student trained in the project is here, not only the two exported above. Each run folder has:
| File | Contents |
|---|---|
student.pt |
The weights. Pixel / latent / flow runs store {"model": EMA state dict, "method", "base"}. SDXS runs store {"trainable": the trained tensors (LoRA, conv_in, VAE, gate, mask), "options"}. All load with torch.load(..., weights_only=True). |
metrics.json |
Final PSNR / LPIPS against the teacher, per edit, training speed and memory |
curve.json |
The learning curve: training loss and periodic validation |
samples.png |
Held-out examples: source, teacher, student |
Closeness to the teacher on held-out people (256×208; the "gain" is PSNR above simply copying the source, which scores about 22.5–23.1 dB):
| Run | Kind | Trained / inference params | Training | Gain (dB) | LPIPS ↓ | Notes |
|---|---|---|---|---|---|---|
kd_test/pixel |
pixel U-Net | 29.4M | 30 min L4, 467 pairs | 2.84 | 0.055 | first test |
kd_test/latent |
latent U-Net (Qwen VAE space) | 33.4M (+127M VAE to decode) | 30 min L4 | 2.64 | 0.111 | blurry |
kd_test/flow |
8-step flow U-Net on the teacher trajectory | 33.5M (+127M VAE) | 30 min L4 | 1.62 | 0.120 | noisy |
pixel_v2/v2a |
pixel U-Net | 29.4M | 90 min L4, 1,098 pairs | 2.94 | 0.061 | = pixel/ |
pixel_v2/v2b |
pixel U-Net + PatchGAN | 29.4M | 90 min L4 | 2.38 | 0.061 | GAN made age worse |
sdxs_v1/turbo_lora |
SDXS, x₀ head, LoRA | 15.9M / 344M | 100 min L40S | 1.78 | 0.081 | |
sdxs_v1/gated_lora |
SDXS, gated head, LoRA, full residual | 15.9M / 344M | 100 min L40S | 2.98 | 0.057 | ghosting on removals |
sdxs_v1/turbo_full |
SDXS, x₀ head, full U-Net | 331M / 331M | 100 min L40S | 3.10 | 0.060 | ghosting |
sdxs_v1/v2_gated_lora_none |
SDXS, gated, LoRA, no residual | 15.9M / 344M | 100 min L40S | 3.19 | 0.056 | softer background |
sdxs_v1/v2_gated_lora_mask |
SDXS, gated, LoRA, learned mask | 15.9M / 344M (331M merged) | 100 min L40S | 3.66 | 0.051 | = sdxs/ |
sdxs_v1/v2_gated_full_mask |
SDXS, gated, full U-Net, learned mask | 331M / 331M | 100 min L40S | 3.66 | 0.052 | ties the LoRA version at 20× the file size |
To run an SDXS checkpoint other than the exported one, or the latent / flow students, use the training code in the
FaceBench repository (src/distill/):
distill.sdxs_export.build_student("student.pt")rebuilds an SDXS student from the base weights;python -m distill.sdxs_export student.pt out_direxports it in the same format assdxs/.
Limitations
- Faces only, CelebA-like crops only. The models were trained on CelebA aligned faces at low resolution.
- Eight fixed edits. Each edit is binary, with no control over strength.
- Inherits the teacher's biases. For example, "older" tends to add grey hair and wrinkles in a fixed style.
- Small evaluation. The human evaluation has 2 graders and 40 requests, so "not significantly different from the teacher" is not proof of equivalence.
- Research use only. The training images are CelebA (non-commercial research licence). The base model is IDKiro/sdxs-512-0.9, and the teacher is Qwen-Image-Edit; check their licences before any other use. Do not use the models to create misleading images of real people.
Related
The teacher outputs and grading data are in the dataset repo
oloruntobiolutola/facebench-pilot-r200.
Its b40u_* folders hold the 40-person human-evaluation set for all six models.
- Downloads last month
- -
Model tree for oloruntobiolutola/facebench-students
Base model
IDKiro/sdxs-512-0.9