RECON checkpoints

Checkpoints trained with the RECON code (saeidrazavi/RECON-DEV-V2). RECON trains diffusion transformers with REPA-style representation alignment and semantic tokens: at block encoder_depth the model predicts the features of a frozen vision encoder (the teacher), pools them into K tokens and appends them to the sequence for the remaining blocks. Sampling pools the tokens from the model's own prediction, so no teacher is needed at inference.

File Task Model Data Steps
c2i/ep-00080.pt Class-to-image SiT-XL/2 ImageNet 256×256 400,320 (80 epochs)
t2i-mscoco/0150000.pt Text-to-image MMDiT MS-COCO 256×256 150,000

Each file is a full training checkpoint with the keys model, ema, opt, args and steps. The evaluation scripts use the EMA weights and read the model configuration from args, so no configuration flags are needed. Training can resume from these files.

Results

Class-to-image (c2i/ep-00080.pt): 50k samples, SDE sampler with 250 steps, no guidance, --seed 0, ADM reference batch.

FID IS Precision Recall
3.2661 206.43 0.8215 0.5556

Text-to-image (t2i-mscoco/0150000.pt): 40,192 MS-COCO val captions, ODE sampler (Heun) with 50 steps, U-ViT FID reference.

CFG scale FID CLIP score
1.0 8.2671 0.2311
2.0 4.5077 0.2465

Training settings

Setting Class-to-image Text-to-image
Model SiT-XL/2 MMDiT
VAE (latents) sdvae-ft-mse-f8d4 sdvae-ft-ema-f8d4
Text condition — CLIP ViT-L/14 features (U-ViT preprocessing)
Teacher dinov2-vit-b dinov2-vit-b
Batch size / learning rate 256 / 1e-4 256 / 1e-4
Mixed precision bf16 fp16
EMA decay 0.9999 0.9999
--proj-coeff 1.5 1.5
--encoder-depth 8 8
Semantic tokens K=1, zeropad head, visible for t ≤ 0.5 K=1, zeropad head, visible for t ≤ 0.5
Caption dropout (--cfg-prob) — 0.1

Usage

Set up the code and data as in the code repository's README. Then download the checkpoints (this repository is private, so run hf auth login first):

uv run hf download SaeedRazavi/RECON-DEV-V2 c2i/ep-00080.pt t2i-mscoco/0150000.pt --local-dir checkpoints/recon

Evaluate them:

# Class-to-image
REFERENCE_NPZ=data/imagenet/VIRTUAL_imagenet256_labeled.npz \
uv run accelerate launch --num_processes 4 --mixed_precision bf16 c2i_eval.py \
    --ckpt checkpoints/recon/c2i/ep-00080.pt

# Text-to-image at CFG 2.0 (reads data/mscoco256 as in training; add --cfg-scale 1.0 for no guidance)
uv run accelerate launch --num_processes 4 t2i_eval.py --ckpt checkpoints/recon/t2i-mscoco/0150000.pt
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support