# Frechet-Distributional-Decoder-Alignment
Decoder weights of FDDA. FDDA post-trains the tokenizer decoder of a frozen latent generative model
with a Fréchet distance loss on generated samples. Code, installation and evaluation:
[github.com/sunset-clouds/Frechet-Distributional-Decoder-Alignment](https://github.com/sunset-clouds/Frechet-Distributional-Decoder-Alignment).
Each checkpoint stores the tokenizer state dict with the aligned decoder (`{"model", "epoch"}`) and
is used with the generator in `--generator_name`. ImageNet 256x256, 50,000 generated samples.
gFDr6 is the mean over six representation spaces (Inception-v3, ConvNeXt-v2, DINOv2, MAE,
SigLIP2, CLIP) of FD divided by the FD of the ImageNet validation set.
## Before generator-side FD post-training
`before_generator_side//`: decoders aligned to the pretrained generators.
Model | `--generator_name` | gFID | IS | gFDr6 | checkpoint
--- | --- |:---:|:---:|:---:| ---
LlamaGen-B | `llamagen-B_256` | 2.74 | 209.3 | 10.88 | [`llamagen-B_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/llamagen/llamagen-B_256.pt)
LlamaGen-L | `llamagen-L_256` | 1.87 | 303.5 | 5.60 | [`llamagen-L_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/llamagen/llamagen-L_256.pt)
GigaTok-S-S | `gigatok-B_256` | 1.68 | 281.6 | 6.59 | [`gigatok-B_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/gigatok/gigatok-B_256.pt)
TiTok-L-32 | `titok-L32` | 1.34 | 202.6 | 6.19 | [`titok-L32.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/titok/titok-L32.pt)
TiTok-B-64 | `titok-B64` | 1.61 | 214.9 | 6.99 | [`titok-B64.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/titok/titok-B64.pt)
VAR-d16 | `var-d16_256` | 1.62 | 280.1 | 7.03 | [`var-d16_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/var/var-d16_256.pt)
VAR-d20 | `var-d20_256` | 1.32 | 296.9 | 5.34 | [`var-d20_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/var/var-d20_256.pt)
VAR-d24 | `var-d24_256` | 1.24 | 304.4 | 4.26 | [`var-d24_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/var/var-d24_256.pt)
iMF-B | `imf-B_256` | 1.65 | 268.1 | 8.54 | [`imf-B_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/imf/imf-B_256.pt)
iMF-L | `imf-L_256` | 1.26 | 281.0 | 5.48 | [`imf-L_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/imf/imf-L_256.pt)
iMF-XL | `imf-XL_256` | 1.14 | 288.3 | 5.03 | [`imf-XL_256.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/before_generator_side/imf/imf-XL_256.pt)
## After generator-side FD post-training
`after_generator_side//`: decoders aligned to the FD post-trained generators (FDAR for
LlamaGen, GigaTok, TiTok and VAR; FD-SIM for iMF).
Model | `--generator_name` | gFID | IS | gFDr6 | checkpoint
--- | --- |:---:|:---:|:---:| ---
LlamaGen-B | `llamagen-B_256-fdpt` | 2.23 | 275.9 | 5.27 | [`llamagen-B_256-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/llamagen/llamagen-B_256-fdpt.pt)
LlamaGen-L | `llamagen-L_256-fdpt` | 1.34 | 313.5 | 3.18 | [`llamagen-L_256-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/llamagen/llamagen-L_256-fdpt.pt)
GigaTok-S-S | `gigatok-B_256-fdpt` | 1.74 | 296.5 | 3.84 | [`gigatok-B_256-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/gigatok/gigatok-B_256-fdpt.pt)
TiTok-L-32 | `titok-L32-fdpt` | 1.35 | 215.1 | 5.49 | [`titok-L32-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/titok/titok-L32-fdpt.pt)
TiTok-B-64 | `titok-B64-fdpt` | 1.42 | 240.3 | 5.42 | [`titok-B64-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/titok/titok-B64-fdpt.pt)
VAR-d16 | `var-d16_256-fdpt` | 1.35 | 307.3 | 2.62 | [`var-d16_256-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/var/var-d16_256-fdpt.pt)
VAR-d20 | `var-d20_256-fdpt` | 1.08 | 308.1 | 1.98 | [`var-d20_256-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/var/var-d20_256-fdpt.pt)
VAR-d24 | `var-d24_256-fdpt` | 1.09 | 308.3 | 1.67 | [`var-d24_256-fdpt.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/var/var-d24_256-fdpt.pt)
iMF-B | `imf-B_256-fdsim` | 0.91 | 304.0 | 4.54 | [`imf-B_256-fdsim.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/imf/imf-B_256-fdsim.pt)
iMF-L | `imf-L_256-fdsim` | 0.80 | 299.1 | 2.43 | [`imf-L_256-fdsim.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/imf/imf-L_256-fdsim.pt)
iMF-XL | `imf-XL_256-fdsim` | 0.78 | 303.7 | 2.21 | [`imf-XL_256-fdsim.pt`](https://huggingface.co/jiajunzhu/Frechet-Distributional-Decoder-Alignment/blob/main/after_generator_side/imf/imf-XL_256-fdsim.pt)
## Reference statistics
`reference_stats/reconstruction/fid_stats/`: ImageNet-validation statistics used by the
reconstruction metrics.
## Download
From the root of the code repository:
```bash
hf download jiajunzhu/Frechet-Distributional-Decoder-Alignment \
--include "*_generator_side/*" --local-dir checkpoints/fdda
hf download jiajunzhu/Frechet-Distributional-Decoder-Alignment \
--include "reference_stats/*" --local-dir .
```
## BibTeX
```bibtex
```