LDM-is-AE: Latent Diffusion is an Intrinsic Auto-Encoder
Released inference checkpoints for LDM-is-AE (NeurIPS 2026). Code: https://github.com/PolyU-VCLab/LDMisAE
Latent diffusion is usually a two-stage pipeline -- train a VAE, then train a diffusion model in its latent space. LDM-is-AE removes the pipeline: the DiT backbone is split into DiT-E / DiT-D and the intermediate feature is supervised in the image domain at every timestep, so the backbone itself becomes an intrinsic auto-encoder trained end-to-end in a single stage.
Checkpoints
| File | Resolution | Keys | Size |
|---|---|---|---|
LDMisAE.256.ckpt |
256x256 | model_ema1, args, epoch |
3.90 GB |
LDMisAE.512.ckpt |
512x512 | model_ema1, args, epoch |
3.93 GB |
Class-conditional ImageNet models. The files are ema1-only (inference); resuming training needs a full checkpoint.
Results
| Setting | FID | IS |
|---|---|---|
| ImageNet 256x256, class-conditional | 1.80 | 314 |
| ImageNet 512x512, class-conditional | 1.90 | 320 |
| Reconstruction (gFID) | 1.82 | -- |
Usage
git clone https://github.com/PolyU-VCLab/LDMisAE.git && cd LDMisAE
pip install -r requirements.txt
huggingface-cli download xtudbxk/LDMisAE LDMisAE.256.ckpt --local-dir weights
CKPT=weights/LDMisAE.256.ckpt IMG_SIZE=256 CFG=2.25 NUM_IMAGES=50000 bash scripts/inference.sh
Keep --ema_mode ema1 (the default); none / ema2 fail fast with a clear error. Sampling needs no
LPIPS weights: --lpips_weight defaults to 0 on the sampling path (LPIPS only enters the training loss).
Citation
@inproceedings{ldm_is_ae_2026,
title = {LDM-is-AE: Latent Diffusion is an Intrinsic Auto-Encoder for End-to-End Image Generation},
author = {Zhang, Zhengqiang and Sun, Lingchen and Wu, Rongyuan and Yi, Qiaosi and
Kong, Xiangtao and Xiao, Chaodong and Zhang, Lei},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}
License
Code is Apache-2.0 (https://github.com/PolyU-VCLab/LDMisAE). Weights are released for research use.