YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Agate 002 โ€” 4-step distilled (unofficial)

An unofficial distillation of LogoLabs/agate-preview-002 (MIT) into a 4-step uniform-grid Euler velocity model with classifier-free guidance baked in (4 network evaluations per image vs the teacher's 100). An 8-step variant is included under 8step/. It drops straight into the Agate webgpu page with steps=4, cfg=1.

from agate import AgatePipeline
pipe = AgatePipeline.from_pretrained("ysharma/agate-002-4step")
images = pipe("a red cube on top of a blue sphere", steps=4, cfg=1.0)   # one image, no CFG

Method: guidance distillation (teacher CFG-3 velocity target, frozen Ettin text encoder, AdamW 1e-5, EMA 0.999) then progressive distillation 32 -> 16 -> 8 -> 4 (Salimans & Ho 2022; Meng et al. 2023), trained on 150k SD-VAE latents from LucasFang/FLUX-Reason-6M (Apache-2.0) with Agate's caption mix (80% short / 20% detail).

Results (256 px, 1 image/prompt, watermark off)

config steps net evals GenEval CLIP FID-5k FD-DINOv2 LPIPS-vs-t50 s/img
teacher50 50 100 0.562 27.03 11.40 130.78 - 0.591
teacher16 16 32 0.569 26.74 11.90 129.98 0.196 0.189
teacher8 8 16 0.533 26.67 13.47 141.18 0.298 0.095
teacher4 4 8 0.466 26.62 20.07 218.23 0.416 0.047
G @50 (GD student) 50 50 0.507 26.39 11.89 131.86 0.218 0.580
s8 (8-step student) 8 8 0.521 26.49 12.35 134.64 0.262 0.093
s4 (4-step student) 4 4 0.527 26.25 12.96 146.47 0.307 0.046

GenEval = official scorer (Mask2Former-Swin-S + OpenCLIP ViT-L/14), mean of the 6 task accuracies; the teacher's published 0.577 uses the official 4-image protocol (0.562 here with 1 image). FID-5k / FD-DINOv2 are against the 5,000 held-out FLUX-Reason images at 256 px and are only comparable within this table. CLIPScore = openai/clip-vit-large-patch14, x100 scale. LPIPS (AlexNet) = vs teacher50 on the first 500 GenEval prompts, same seed. Latency = batch-1 a100-large with CUDA graphs.

At the same 8 network evaluations, the 4-step student beats the teacher run at 4 steps on GenEval (0.527 vs 0.466), FID-5k (12.96 vs 20.07) and FD-DINOv2 (146.5 vs 218.2), and lands closer to the teacher's 50-step output (LPIPS 0.307 vs 0.416); it trails it on CLIPScore by 0.37 points (26.25 vs 26.62).

Limitations

  • No negative prompts: CFG is baked in; negative_prompt is ignored at cfg=1.0.
  • 256 px only, like the teacher.
  • Inherits the teacher's weaknesses (counting above ~3, negation, exact text) and its unprompted-nudity tendency โ€” use an output classifier before public deployment.
  • Guidance-distillation gate (LPIPS G@50 vs teacher@50-CFG) measured 0.174, above the 0.15 target; stage LPIPS vs each stage's teacher: s16 0.015, s8 0.032, s4 0.104.
  • The watermark (payload AGATE002) is on by default; pass watermark=False for images that go into metrics.

Credits

Model and code: LogoLabs (Logolabs/agate-preview-002, MIT). Unofficial distillation by ysharma; distilled weights under the same MIT license as the teacher's code.

Images

The same 16 prompts and seeds, left to right โ€” teacher (002) at 50 steps with CFG 3, the 8-step student, the 4-step student:

teacher @50 (CFG 3) s8 @8 steps s4 @4 steps

Training

Per-stage distillation loss (log scale), MSE against each stage's teacher target:

Stage checkpoints and logs: ysharma/agate-002-distill-artifacts (g/, s16/, s8/, s4/, logs/, grids/). Live dashboards: agate-002-distill-trackio. Evaluation data and raw results: ysharma/agate-002-eval-images.

Stage LPIPS against each stage's teacher (8 fixed prompts, same noise): s16 0.015, s8 0.032, s4 0.104. Guidance-distillation gate: G@50 vs teacher@50-CFG 0.174 (above the 0.15 target โ€” recorded as a finding); G@32 vs G@50 0.004.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support