data-archetype/dinac3_96
dinac3_96 is a deterministic semantic autoencoder for latent diffusion. It maps an RGB image to a 96-channel latent grid at stride 16 and decodes it back in a single forward pass.
It is strongly inspired by semantic_vae, the latent with the fastest downstream DiT convergence we know of. dinac3_96 aims to keep that convergence while improving reconstruction and simplifying the encoder to a single DINOv3 pass instead of two.
- Encoder: a frozen DINOv3 ViT-B/16 behind a trainable patch embedding. All twelve blocks are read, standardized per channel and mapped to the latent by one trainable linear projection.
- Latent: 32 semantic channels, held during training to a fixed random projection of DINOv3-B's blocks summed with fitted block weights, plus 64 reconstruction channels (reconstruction losses and VISReg). No posterior noise.
- Decoder: a 4-block transformer trunk at width 1152 followed by a dense convolutional up-path, a slightly modernized VQGAN-like decoder (stride 16 to full resolution, three residual blocks per level).
- Training losses: mainly a DINOv3-B feature loss on the reconstruction, with pixel and blurred-image MSE terms each at about a tenth of its gradient, following semantic_vae's DINO-loss-trained decoder.
Technical report · Reconstruction gallery (39 images: original, reconstruction, RGB difference, latent PCA, per-image PSNR)
Generation benchmark
We evaluate a latent by training a class-conditional DiT on it (about 190M parameters, 100k steps at batch 256 on ImageNet-1k, flow matching with SPRINT routing). We score 10,000 samples against 810,000 real images on our mixed-resolution ImageNet benchmark: four aspect ratios at 256- and 384-px areas. Lower is better.[^fid] For each model: 50 NFEs, PDG 2.5–4.0, best settings kept. Bold marks the better value in each column. The DiT was trained on dinac3_96's latent (the released latent space, frozen since the joint phase), and its samples were decoded by the released decoder.
| model | FID | MIND | Monge-DINO |
|---|---|---|---|
| dinac3_96 | 9.04 |
8.29 |
10.92 |
| semantic_vae | 10.02 |
7.42 |
15.94 |
Monge-DINO uses DINOv3-B features, the same network whose features the decoder is trained to match, so it is not independent of the training loss; FID and MIND, on Inception features, are. The same holds for rMonge-DINOv3-B below.
Our blind pairwise comparison of the two models' generations put them on equal footing.
Protocol, DiT settings and caveats: technical report, section 9.1.
Reconstruction
10,000 ImageNet images: 10 real images per class with the generation benchmark's aspect-ratio quotas (9 at 256-px area and 1 at 384 px per class). Metrics compare the reconstructions with the same 10,000 originals.[^rfid] semantic_vae runs at its native bf16.
| model | rFID | rMIND | rMonge-DINOv3-B | PSNR mean | SSIM | LPIPS-VGG | LPIPS-Alex |
|---|---|---|---|---|---|---|---|
| dinac3_96 | 0.809 |
0.166 |
1.205 |
28.16 |
0.8195 |
0.0970 |
0.0362 |
| semantic_vae | 2.149 |
0.820 |
5.994 |
25.19 |
0.7103 |
0.1715 |
0.0669 |
PSNR here: reconstructions clamped and rounded to uint8, peak 255, mean of per-image values.
2k PSNR benchmark (the image set of our earlier releases):
| Model | Mean PSNR (dB) | Std (dB) | Median (dB) | P5 (dB) | P95 (dB) |
|---|---|---|---|---|---|
| dinac3_96 | 32.15 |
5.11 |
31.64 |
24.27 |
40.70 |
| semantic_vae | 28.03 |
4.58 |
27.79 |
20.98 |
35.44 |
| dinac_ae_d2 | 35.59 |
4.87 |
35.40 |
27.89 |
43.51 |
| FLUX.2 VAE | 36.28 |
4.53 |
36.07 |
28.89 |
43.63 |
Latent interface
- 96 channels at stride 16; image height and width must be multiples of 16.
- Channels 0–63 are reconstruction channels, 64–95 semantic channels
(
free_channels(z),semantic_channels(z)). encode(images): RGB in [-1, 1] → deterministic FP32 latents, whitened per channel with the shipped statistics (what our DiTs were trained on).decode(latents, height, width): whitened latents → FP32 RGB in about [-1, 1], unclamped. Any multiple of 16 works, including sizes above 1024 px.encode_raw/decode_rawuse the unwhitened latent;whiten/dewhitenconvert.
Precision
- Set the dtype only through
from_pretrained(dtype=...); dtype casts after loading raise aTypeError(.to(device)works). torch.bfloat16(default) runs under bf16 autocast, bit-identical to the training checkpoint under bf16 autocast.torch.float32runs without autocast.- Encoder and decoder are compiled by default (about 20 s per module on the
first call). Pass
compile_encoder=False, compile_decoder=Falsefor eager.
Speed
RTX 5090, bf16, compiled (the default); mean of 20 timed batches after 3 warm-up batches. Peak VRAM includes the weights.
| Resolution | Batch | Encode ms/image | Decode ms/image | Peak VRAM encode / decode (GiB) |
|---|---|---|---|---|
256x256 |
128 |
0.41 |
1.51 |
3.3 / 8.5 |
512x512 |
32 |
1.79 |
6.10 |
3.3 / 8.5 |
1024x1024 |
8 |
9.03 |
25.4 |
3.3 / 8.5 |
2048x2048 |
1 |
69.4 |
157 |
1.8 / 6.4 |
Eager mode decodes about 2.4× slower (1.8× at 2048²) and peaks at 12.7 GiB on the batched rows.
Size limit. No resolution limit is coded; GPU memory is the practical one. On a 32 GB RTX 5090, compiled decoding handles a single 4096 × 4096 image (15.8 s, 24.5 GiB peak); eager decoding runs out of memory at that size.
Usage
Download the repository and install the package, in an environment with a
CUDA-enabled PyTorch (dependencies: torch>=2.13, timm==1.0.26, safetensors,
huggingface-hub):
hf download data-archetype/dinac3_96 --local-dir dinac3_96
pip install ./dinac3_96
import torch
from dinac3 import Dinac3
model = Dinac3.from_pretrained(
"data-archetype/dinac3_96", # or a local directory
device="cuda",
dtype=torch.bfloat16,
)
images = ... # [B, 3, H, W] on CUDA, RGB in [-1, 1], H and W multiples of 16
latents = model.encode(images) # [B, 96, H/16, W/16], whitened
semantic = model.semantic_channels(latents) # [B, 32, H/16, W/16]
recon = model.decode(latents, images.shape[-2], images.shape[-1])
recon = recon.clamp(-1, 1) # only for display or saving
from_pretrained(path_or_repo, *, device, dtype=torch.bfloat16, compile_encoder=True, compile_decoder=True, revision=None):
- A
Pathis always a local directory. - A
stris a local directory if one exists. - A
strthat looks like a path (it starts with.,/or~, or has more than one/) but names no directory raises an error. - Any other
stris a Hub repository id; onlyconfig.jsonandmodel.safetensorsare downloaded. - The DINOv3-B weights ship inside
model.safetensors, so a local copy loads offline. - Inference is CUDA only.
Details
- Network: 164M parameters, 79M of them trained.
- Encoder (87M): frozen DINOv3 ViT-B/16 behind a trainable patch embedding; all 12 blocks standardized, concatenated and projected linearly to 96 channels at stride 16.
- Decoder (78M): 1×1 projection to width 1152, 4 transformer blocks, then a convolutional up-path from stride 16 to full resolution (512-channel handoff, levels of 256, 256, 128 and 128 channels, three residual 3 × 3 blocks per level).
- Training data: about 14 million images from a mix of public and licensed or curated datasets: mostly photographs, plus book covers, a few text-heavy datasets and 1% synthetic rendered text. Aspect-ratio buckets, downsampled only. No training images are redistributed with the model.
- Training:
- Losses: mainly a DINOv3-B all-block feature MSE on the reconstruction, plus pixel and blurred-image MSE at about a tenth of its gradient; while the encoder trained, also a semantic alignment loss and VISReg.
- Steps: about 150k at 256/384 px (batch 128) plus about 25k fine-tuning at five resolutions from 256 to 1024 px (batch 32). The free encoder weights (patch embedding and output projection) were trained jointly with the decoder for the first 52k steps and then frozen.
- Matrix weights trained with Dion (dionw), the rest with AdamW. The released weights are EMA weights.
- More: the technical report covers the semantic target, the decoder heads tried, loss weights, training phases, ablations and the benchmark protocol.
- Related: semantic_vae, DINAC-AE-D2.
License
- Code (the
dinac3package and scripts): Copyright 2026 data-archetype, licensed under the Apache License 2.0, see LICENSE. - Weights: the DINOv3 License, see LICENSE-DINOV3.md.
The weights contain DINOv3-B's weights and a trained derivative of its patch
embedding, so under the DINOv3 License (section 1.b.i) they may only be
distributed under that licence, which must accompany them. Its terms include:
- acknowledging the use of DINO Materials in publications (section 1.b.ii);
- no reverse engineering (section 1.b.iv);
- compliance with trade controls, which excludes ITAR-regulated uses and military, nuclear, espionage and weapons applications (sections 1.b.iii and 1.b.v).
The license: other tag above refers to the weights' DINOv3 License. See
ATTRIBUTION.md for sources and design credits.
Citation
@misc{dinac3_96,
title = {dinac3_96: a DINOv3-based semantic autoencoder with a one-pass DINO-loss decoder},
author = {data-archetype},
email = {data-archetype@proton.me},
year = {2026},
month = oct,
url = {https://huggingface.co/data-archetype/dinac3_96},
}
[^fid]: Papers usually report FID on 50,000 square samples at a single resolution, so our numbers are not directly comparable with published ones.
[^rfid]: Papers usually report rFID on the 50,000 ImageNet validation images at a single square resolution, so our numbers are not directly comparable with published ones.
- Downloads last month
- -
Model tree for data-archetype/dinac3_96
Base model
timm/vit_base_patch16_dinov3.lvd1689m