LatentMaid VAE

An image autoencoder packaged as a diffusers model. Its API is the same as AutoencoderKL, so it can stand in for any diffusers VAE.

Downsampling 8x (256x256 image -> 32x32 latent)
Latent channels 4
Input / output RGB in [-1, 1]
Parameters 54.4M

Usage

The model class ships in this repo (modeling_latentmaidvae.py), so load it with trust_remote_code=True:

import torch
from diffusers import AutoModel

vae = AutoModel.from_pretrained("KBlueLeaf/latentmaid-vae", trust_remote_code=True).eval()

x = torch.rand(1, 3, 256, 256) * 2 - 1                  # image in [-1, 1]
with torch.no_grad():
    posterior = vae.encode(x).latent_dist              # DiagonalGaussianDistribution
    z = posterior.mode()                               # [1, 4, 32, 32]
    x_rec = vae.decode(z).sample                       # [1, 3, 256, 256]

encode returns AutoencoderKLOutput and decode returns DecoderOutput, the same as AutoencoderKL.

Latent statistics

config.json stores per-channel latent statistics for normalising latents before using them in a generative model:

mean = torch.tensor(vae.config.latents_mean).view(1, -1, 1, 1)
std = torch.tensor(vae.config.latents_std).view(1, -1, 1, 1)
z_norm = (z - mean) / std
z = z_norm * std + mean                                # before vae.decode

scaling_factor is 1.0; use the per-channel statistics above instead.

Downloads last month
7
Safetensors
Model size
54.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support