Agate Preview 002, 4-step (unofficial distillation)

An unofficial distillation of LogoLabs' Agate Preview 002 (260M parameters, MIT) into a 4-step model with classifier-free guidance baked in: 4 network passes per image instead of the teacher's 100 (50 Euler steps × a conditional and an unconditional pass). Same architecture, same 256 px output, same agate/ code; only the weights differ.

Try it: live, redraws on every keystroke (GPU) · in your browser (WebGPU, nothing leaves your machine) · everything in one collection

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("ML-Intern-lab/agate-preview-002-4step")
sys.path.insert(0, path)
from agate import AgatePipeline

pipe = AgatePipeline.from_pretrained(path, device="cuda")
image = pipe("a red cube on top of a blue sphere", steps=4, cfg=1.0, seed=0)[0]   # cfg=1.0: one pass per step
image.save("cube.png")

Use steps=4, cfg=1.0. The bundled pipeline skips the unconditional pass when cfg == 1.0. Every image carries Agate's invisible watermark (payload AGATE002) and provenance metadata; pass watermark=False only for images that go into metrics. A WebGPU build (webgpu/, same graph I/O as Agate's own) runs in the browser with onnxruntime-web.

Results

GenEval with the official scorer (Mask2Former + CLIP), 4 images per prompt (2,212 images per row), mean of the six task accuracies, with 95% bootstrap intervals over the 553 prompts. Image metrics on 5,000 held-out FLUX-Reason captions at 256 px.

model network passes GenEval 95% CI FID-5k FD-DINOv2 CLIPScore
teacher, 50 steps + CFG 3 100 0.563 [0.534, 0.591] 11.40 130.8 27.03
teacher, 16 steps + CFG 3 32 0.564 [0.535, 0.592] 11.90 130.0 26.74
this model (s4p, 4 steps) 4 0.536 [0.508, 0.564] 12.57 142.2 26.97
previous version (s4, 4 steps, in v1/) 4 0.509 [0.480, 0.538] 12.96 146.5 26.25
guidance-distilled model before step halving, 4 steps 4 0.447 [0.418, 0.474] – – –
teacher, 4 steps + CFG 3 (1 image per prompt) 8 0.466 – 20.07 218.2 26.62

Paired differences (same prompts and seeds): this model vs the previous version +0.027 [+0.014, +0.041]; vs the teacher at 50 steps −0.027. Per task (this model / teacher at 50 steps): single object 0.91 / 0.92, two objects 0.56 / 0.63, counting 0.33 / 0.38, colours 0.77 / 0.74, position 0.23 / 0.27, colour attribution 0.41 / 0.45.

Speed: about 0.046 s per image at batch 1 on an A100 with CUDA graphs, against 0.59 s for the teacher at 50 steps (12.8× faster; measured for the previous version, and this version has the identical network and step count). FID and FD-DINOv2 are only comparable within this table.

Comparison grid: 16 GenEval prompts; top row teacher at 50 steps, middle row the previous version, bottom row this model

Top row: teacher at 50 steps with CFG. Middle: previous version (s4). Bottom: this model (s4p). The two student rows share seeds; the teacher row was generated with different seeds, so compare the student rows with each other, and the teacher row only for overall look.

How it was made

Two runs by ML Intern on Hugging Face Jobs, about $37 of GPU in total.

  1. Run 1 (29–30 Sep, $22.43):
    • Guidance distillation: a student learns the teacher's CFG-3 velocity in one pass (3,000 steps).
    • Progressive distillation: the step count is halved 32 → 16 → 8 → 4 (Salimans & Ho 2022; Meng et al. 2023).
    • Data: 150k SD-VAE latents from LucasFang/FLUX-Reason-6M, Agate's own training set (Apache-2.0).
  2. Run 2 (30 Sep – 1 Oct, $14.70): a perceptual fine-tune of the 4-step model.
    • Targets: 24k images from the teacher at 16 steps with CFG, each paired with its starting noise. GenEval and held-out captions were excluded.
    • Training: backpropagation through all 4 sampling steps. Loss = latent MSE + LPIPS + DINOv2 cosine on TAESD decodes; 1,125 steps.
    • Ship gate: shipped only after passing a gate on GenEval, LPIPS and FID.

Per-stage logs, checkpoints, scripts and every evaluation image are kept in private repos by the author; the evaluation images and raw scores are in agate-preview-002-4step-eval (private for now).

Run 1 results (1 image per prompt, for reference)

model steps network passes GenEval FID-5k LPIPS to teacher@50 s / image (A100)
teacher 50 100 0.562 11.40 – 0.591
teacher 16 32 0.569 11.90 0.196 0.189
teacher 8 16 0.533 13.47 0.298 0.095
teacher 4 8 0.466 20.07 0.416 0.047
8-step student (8step/) 8 8 0.521 12.35 0.262 0.093
4-step student (v1/) 4 4 0.527 12.96 0.307 0.046

With one image per prompt each GenEval number is only accurate to about ±0.02; the 4-image table above supersedes it.

Files

path what
generator.safetensors this model (s4p), bf16
v1/generator.safetensors the previous 4-step version (s4, run 1)
8step/generator.safetensors the 8-step student from run 1 (use steps=8, cfg=1.0)
agate/, text_encoder/, config.json Agate 002's code and text encoder; config.json sets 4 steps, CFG 1
webgpu/ ONNX build for the browser (opset 17, fp16 weights, fp32 compute); parity ≤ 2/255 per pixel vs PyTorch

Limitations

  • No negative prompts. Guidance is baked into the weights; negative_prompt is ignored at cfg=1.0.
  • 256 × 256 only, like the teacher.
  • Agate's weaknesses remain: exact text, counts above three, negation. The student is still about 0.03 GenEval below the teacher, mostly on two-object and colour-attribution prompts.
  • People can come out partly unclothed without being asked, a tendency the teacher's card documents. Use an output classifier for public deployments; the live Space does.
  • LPIPS in run 2 was measured on 200 prompts in single-pair processes after batched loops gave inconsistent values; treat LPIPS as supporting evidence and GenEval and FID as the main results.

Credits and licence

Model architecture, code and teacher weights: LogoLabs, Agate Preview 002 (MIT). Training data: FLUX-Reason-6M (Apache-2.0). Distillation by ML Intern for ysharma; not affiliated with LogoLabs. The distilled weights are released under the MIT licence, like the teacher.

Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ML-Intern-lab/agate-preview-002-4step

Finetuned
(1)
this model

Dataset used to train ML-Intern-lab/agate-preview-002-4step