RECON checkpoints
Checkpoints trained with the RECON code (saeidrazavi/RECON-DEV-V2).
RECON trains diffusion transformers with REPA-style representation alignment and semantic tokens:
at block encoder_depth the model predicts the features of a frozen vision encoder (the teacher),
pools them into K tokens and appends them to the sequence for the remaining blocks. Sampling pools
the tokens from the model's own prediction, so no teacher is needed at inference.
| File | Task | Model | Data | Steps |
|---|---|---|---|---|
c2i/ep-00080.pt |
Class-to-image | SiT-XL/2 | ImageNet 256×256 | 400,320 (80 epochs) |
t2i-mscoco/0150000.pt |
Text-to-image | MMDiT | MS-COCO 256×256 | 150,000 |
Each file is a full training checkpoint with the keys model, ema, opt, args and steps.
The evaluation scripts use the EMA weights and read the model configuration from args, so no
configuration flags are needed. Training can resume from these files.
Results
Class-to-image (c2i/ep-00080.pt): 50k samples, SDE sampler with 250 steps, no guidance,
--seed 0, ADM reference batch.
| FID | IS | Precision | Recall |
|---|---|---|---|
| 3.2661 | 206.43 | 0.8215 | 0.5556 |
Text-to-image (t2i-mscoco/0150000.pt): 40,192 MS-COCO val captions, ODE sampler (Heun) with
50 steps, U-ViT FID reference.
| CFG scale | FID | CLIP score |
|---|---|---|
| 1.0 | 8.2671 | 0.2311 |
| 2.0 | 4.5077 | 0.2465 |
Training settings
| Setting | Class-to-image | Text-to-image |
|---|---|---|
| Model | SiT-XL/2 | MMDiT |
| VAE (latents) | sdvae-ft-mse-f8d4 |
sdvae-ft-ema-f8d4 |
| Text condition | — | CLIP ViT-L/14 features (U-ViT preprocessing) |
| Teacher | dinov2-vit-b |
dinov2-vit-b |
| Batch size / learning rate | 256 / 1e-4 | 256 / 1e-4 |
| Mixed precision | bf16 | fp16 |
| EMA decay | 0.9999 | 0.9999 |
--proj-coeff |
1.5 | 1.5 |
--encoder-depth |
8 | 8 |
| Semantic tokens | K=1, zeropad head, visible for t ≤ 0.5 |
K=1, zeropad head, visible for t ≤ 0.5 |
Caption dropout (--cfg-prob) |
— | 0.1 |
Usage
Set up the code and data as in the code repository's README. Then download the checkpoints
(this repository is private, so run hf auth login first):
uv run hf download SaeedRazavi/RECON-DEV-V2 c2i/ep-00080.pt t2i-mscoco/0150000.pt --local-dir checkpoints/recon
Evaluate them:
# Class-to-image
REFERENCE_NPZ=data/imagenet/VIRTUAL_imagenet256_labeled.npz \
uv run accelerate launch --num_processes 4 --mixed_precision bf16 c2i_eval.py \
--ckpt checkpoints/recon/c2i/ep-00080.pt
# Text-to-image at CFG 2.0 (reads data/mscoco256 as in training; add --cfg-scale 1.0 for no guidance)
uv run accelerate launch --num_processes 4 t2i_eval.py --ckpt checkpoints/recon/t2i-mscoco/0150000.pt