FaceBench students: one-step face editors distilled from Qwen-Image-Edit

Two small image-to-image models that perform eight binary face edits in a single forward pass. They were distilled from Qwen-Image-Edit (20B parameters, 50 steps, about 52 s per edit on an H100) as part of the FaceBench project, which benchmarks identity-preserving face editing.

SDXS student: source, then the 8 edits SDXS student on a person not seen in training. From left to right: source, add glasses, remove glasses, add beard, remove beard, smile, neutral, older, younger.

Folder Model Parameters Resolution Speed
sdxs/ SDXS student (recommended): SDXS-512-0.9 with LoRA merged in, gated one-step head, learned copy-back mask 331M 512×416 38 ms one at a time, 11 ms per image batched (L40S, fp16); about 0.3 s on an Apple M-series GPU
pixel/ Pixel student: hybrid compact U-Net trained from scratch, predicts a residual on the source 29M 256×208 29 ms batched (L4)

Edits

Edit name Meaning Teacher prompt it was distilled from
Eyeglasses+ add glasses "Add eyeglasses."
Eyeglasses- remove glasses "Remove the eyeglasses."
No_Beard+ add a beard "Add a beard."
No_Beard- remove the beard "Remove the beard."
Smiling+ smile "Make the person smile."
Smiling- neutral expression "Change the facial expression to neutral, not smiling, with the lips closed."
Young+ older "Make the same person look older, keeping their face shape, gender, hairstyle and background unchanged apart from age."
Young- younger "Make the same person look like a young adult, keeping their face shape, skin tone, ethnicity, gender, hairstyle, accessories and background unchanged apart from age."

The edits are fixed: the prompt embeddings are stored with the model (sdxs/editor.safetensors), so no text encoder is needed. Free-text prompts are not supported.

Usage

pip install torch diffusers safetensors pillow huggingface_hub
hf download oloruntobiolutola/facebench-students --include "inference.py" --local-dir .

python inference.py --model sdxs --edit Eyeglasses+ --input face.jpg --out out/
python inference.py --model sdxs --edit all --input photos/ --out out/ --sheet    # every edit, plus a sheet per photo
python inference.py --model pixel --edit Young+ --input face.jpg --out out/

inference.py downloads the weights it needs from this repo. From Python:

from PIL import Image
from inference import load

ed = load("sdxs", device="cuda")            # or "pixel"; device defaults to cuda > mps > cpu
ed.edit(Image.open("face.jpg"), "Smiling+").save("smile.png")

Inputs should be face crops framed like CelebA aligned photos (face centred, about 178×218). Other photos are resized to the model's resolution, and results degrade the further they are from that framing.

How it works

SDXS student.

  1. The tiny VAE encodes the source.
  2. One U-Net step at t = 999 runs with the edit's fixed prompt embedding, which gives an x₀ estimate.
  3. A learned per-channel gate moves the source latent towards that estimate: z = z_src + g·(x̂₀ − z_src).
  4. The VAE decodes the result.
  5. A learned mask, sigmoid(b − k·blur|D(z) − D(z_src)|), copies the original pixels back wherever the edit left the image unchanged. This keeps backgrounds sharp and removes ghosting on removals.

It was trained with LoRA (rank 32) on the U-Net, plus the VAE and conv_in, using L1 + LPIPS against the teacher's edits.

Pixel student. A 29M U-Net with convolutional ResBlocks at high resolution, self-attention at 32×26 and a FiLM edit embedding. It predicts output = source + f(source, edit), trained with L1 + LPIPS.

Training data. About 2,000 Qwen-Image-Edit edits of CelebA photos, with balanced edit types. Validation used people held out from training.

Evaluation

Closeness to the teacher. Measured at 256×208 on held-out people, against the teacher's own edit:

Model PSNR gain over copying the source LPIPS ↓
SDXS student +3.66 dB 0.051
Pixel student +2.94 dB 0.061

Blind human evaluation. Two graders judged 40 requests (5 per edit type) on 40 people not in the training data. Each request was edited by six models, shown in random order with names hidden, at equal display size. Graders agreed reliably on whether the edit was done (Krippendorff's α = 0.81) and on same person (α = 0.77 ordinal, 0.82 interval).

Model Edit done Done and same person ≥ 4/5 Same person (1–5) Quality (1–5) Picked best
Qwen-Image-Edit (teacher) 96% 91% 4.47 4.69 30%
SDXS student 96% 88% 4.36 4.45 29%
FLUX.1 Kontext 70% 70% 4.74 4.74 20%
LLaDA-Image 86% 51% 3.27 4.20 6%
Pixel student 56% 52% 3.44 3.09 10%
SDXS-512 untrained (zero-shot img2img) 11% 2% 1.39 1.84 1%

The SDXS student is not significantly different from its teacher on any measure (paired tests, Holm-corrected). It is ahead of FLUX.1 Kontext on success and of LLaDA-Image on identity.

The pixel student handles the local additions (glasses, beard, smile) but fails the age edits and glasses removal.

All trained checkpoints (checkpoints/)

Every student trained in the project is here, not only the two exported above. Each run folder has:

File Contents
student.pt The weights. Pixel / latent / flow runs store {"model": EMA state dict, "method", "base"}. SDXS runs store {"trainable": the trained tensors (LoRA, conv_in, VAE, gate, mask), "options"}. All load with torch.load(..., weights_only=True).
metrics.json Final PSNR / LPIPS against the teacher, per edit, training speed and memory
curve.json The learning curve: training loss and periodic validation
samples.png Held-out examples: source, teacher, student

Closeness to the teacher on held-out people (256×208; the "gain" is PSNR above simply copying the source, which scores about 22.5–23.1 dB):

Run Kind Trained / inference params Training Gain (dB) LPIPS ↓ Notes
kd_test/pixel pixel U-Net 29.4M 30 min L4, 467 pairs 2.84 0.055 first test
kd_test/latent latent U-Net (Qwen VAE space) 33.4M (+127M VAE to decode) 30 min L4 2.64 0.111 blurry
kd_test/flow 8-step flow U-Net on the teacher trajectory 33.5M (+127M VAE) 30 min L4 1.62 0.120 noisy
pixel_v2/v2a pixel U-Net 29.4M 90 min L4, 1,098 pairs 2.94 0.061 = pixel/
pixel_v2/v2b pixel U-Net + PatchGAN 29.4M 90 min L4 2.38 0.061 GAN made age worse
sdxs_v1/turbo_lora SDXS, x₀ head, LoRA 15.9M / 344M 100 min L40S 1.78 0.081
sdxs_v1/gated_lora SDXS, gated head, LoRA, full residual 15.9M / 344M 100 min L40S 2.98 0.057 ghosting on removals
sdxs_v1/turbo_full SDXS, x₀ head, full U-Net 331M / 331M 100 min L40S 3.10 0.060 ghosting
sdxs_v1/v2_gated_lora_none SDXS, gated, LoRA, no residual 15.9M / 344M 100 min L40S 3.19 0.056 softer background
sdxs_v1/v2_gated_lora_mask SDXS, gated, LoRA, learned mask 15.9M / 344M (331M merged) 100 min L40S 3.66 0.051 = sdxs/
sdxs_v1/v2_gated_full_mask SDXS, gated, full U-Net, learned mask 331M / 331M 100 min L40S 3.66 0.052 ties the LoRA version at 20× the file size

To run an SDXS checkpoint other than the exported one, or the latent / flow students, use the training code in the FaceBench repository (src/distill/):

  • distill.sdxs_export.build_student("student.pt") rebuilds an SDXS student from the base weights;
  • python -m distill.sdxs_export student.pt out_dir exports it in the same format as sdxs/.

Limitations

  • Faces only, CelebA-like crops only. The models were trained on CelebA aligned faces at low resolution.
  • Eight fixed edits. Each edit is binary, with no control over strength.
  • Inherits the teacher's biases. For example, "older" tends to add grey hair and wrinkles in a fixed style.
  • Small evaluation. The human evaluation has 2 graders and 40 requests, so "not significantly different from the teacher" is not proof of equivalence.
  • Research use only. The training images are CelebA (non-commercial research licence). The base model is IDKiro/sdxs-512-0.9, and the teacher is Qwen-Image-Edit; check their licences before any other use. Do not use the models to create misleading images of real people.

Related

The teacher outputs and grading data are in the dataset repo oloruntobiolutola/facebench-pilot-r200. Its b40u_* folders hold the 40-person human-evaluation set for all six models.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oloruntobiolutola/facebench-students

Finetuned
(3)
this model