data-archetype/dinac3_96

dinac3_96 is a deterministic semantic autoencoder for latent diffusion. It maps an RGB image to a 96-channel latent grid at stride 16 and decodes it back in a single forward pass.

It is strongly inspired by semantic_vae, the latent with the fastest downstream DiT convergence we know of. dinac3_96 aims to keep that convergence while improving reconstruction and simplifying the encoder to a single DINOv3 pass instead of two.

  • Encoder: a frozen DINOv3 ViT-B/16 behind a trainable patch embedding. All twelve blocks are read, standardized per channel and mapped to the latent by one trainable linear projection.
  • Latent: 32 semantic channels, held during training to a fixed random projection of DINOv3-B's blocks summed with fitted block weights, plus 64 reconstruction channels (reconstruction losses and VISReg). No posterior noise.
  • Decoder: a 4-block transformer trunk at width 1152 followed by a dense convolutional up-path, a slightly modernized VQGAN-like decoder (stride 16 to full resolution, three residual blocks per level).
  • Training losses: mainly a DINOv3-B feature loss on the reconstruction, with pixel and blurred-image MSE terms each at about a tenth of its gradient, following semantic_vae's DINO-loss-trained decoder.

Technical report · Reconstruction gallery (39 images: original, reconstruction, RGB difference, latent PCA, per-image PSNR)

Generation benchmark

We evaluate a latent by training a class-conditional DiT on it (about 190M parameters, 100k steps at batch 256 on ImageNet-1k, flow matching with SPRINT routing). We score 10,000 samples against 810,000 real images on our mixed-resolution ImageNet benchmark: four aspect ratios at 256- and 384-px areas. Lower is better.[^fid] For each model: 50 NFEs, PDG 2.5–4.0, best settings kept. Bold marks the better value in each column. The DiT was trained on dinac3_96's latent (the released latent space, frozen since the joint phase), and its samples were decoded by the released decoder.

model FID MIND Monge-DINO
dinac3_96 9.04 8.29 10.92
semantic_vae 10.02 7.42 15.94

Monge-DINO uses DINOv3-B features, the same network whose features the decoder is trained to match, so it is not independent of the training loss; FID and MIND, on Inception features, are. The same holds for rMonge-DINOv3-B below.

Our blind pairwise comparison of the two models' generations put them on equal footing.

Protocol, DiT settings and caveats: technical report, section 9.1.

Reconstruction

10,000 ImageNet images: 10 real images per class with the generation benchmark's aspect-ratio quotas (9 at 256-px area and 1 at 384 px per class). Metrics compare the reconstructions with the same 10,000 originals.[^rfid] semantic_vae runs at its native bf16.

model rFID rMIND rMonge-DINOv3-B PSNR mean SSIM LPIPS-VGG LPIPS-Alex
dinac3_96 0.809 0.166 1.205 28.16 0.8195 0.0970 0.0362
semantic_vae 2.149 0.820 5.994 25.19 0.7103 0.1715 0.0669

PSNR here: reconstructions clamped and rounded to uint8, peak 255, mean of per-image values.

2k PSNR benchmark (the image set of our earlier releases):

Model Mean PSNR (dB) Std (dB) Median (dB) P5 (dB) P95 (dB)
dinac3_96 32.15 5.11 31.64 24.27 40.70
semantic_vae 28.03 4.58 27.79 20.98 35.44
dinac_ae_d2 35.59 4.87 35.40 27.89 43.51
FLUX.2 VAE 36.28 4.53 36.07 28.89 43.63

Latent interface

  • 96 channels at stride 16; image height and width must be multiples of 16.
  • Channels 0–63 are reconstruction channels, 64–95 semantic channels (free_channels(z), semantic_channels(z)).
  • encode(images): RGB in [-1, 1] → deterministic FP32 latents, whitened per channel with the shipped statistics (what our DiTs were trained on).
  • decode(latents, height, width): whitened latents → FP32 RGB in about [-1, 1], unclamped. Any multiple of 16 works, including sizes above 1024 px.
  • encode_raw / decode_raw use the unwhitened latent; whiten / dewhiten convert.

Precision

  • Set the dtype only through from_pretrained(dtype=...); dtype casts after loading raise a TypeError (.to(device) works).
  • torch.bfloat16 (default) runs under bf16 autocast, bit-identical to the training checkpoint under bf16 autocast. torch.float32 runs without autocast.
  • Encoder and decoder are compiled by default (about 20 s per module on the first call). Pass compile_encoder=False, compile_decoder=False for eager.

Speed

RTX 5090, bf16, compiled (the default); mean of 20 timed batches after 3 warm-up batches. Peak VRAM includes the weights.

Resolution Batch Encode ms/image Decode ms/image Peak VRAM encode / decode (GiB)
256x256 128 0.41 1.51 3.3 / 8.5
512x512 32 1.79 6.10 3.3 / 8.5
1024x1024 8 9.03 25.4 3.3 / 8.5
2048x2048 1 69.4 157 1.8 / 6.4

Eager mode decodes about 2.4× slower (1.8× at 2048²) and peaks at 12.7 GiB on the batched rows.

Size limit. No resolution limit is coded; GPU memory is the practical one. On a 32 GB RTX 5090, compiled decoding handles a single 4096 × 4096 image (15.8 s, 24.5 GiB peak); eager decoding runs out of memory at that size.

Usage

Download the repository and install the package, in an environment with a CUDA-enabled PyTorch (dependencies: torch>=2.13, timm==1.0.26, safetensors, huggingface-hub):

hf download data-archetype/dinac3_96 --local-dir dinac3_96
pip install ./dinac3_96
import torch
from dinac3 import Dinac3

model = Dinac3.from_pretrained(
    "data-archetype/dinac3_96",  # or a local directory
    device="cuda",
    dtype=torch.bfloat16,
)

images = ...  # [B, 3, H, W] on CUDA, RGB in [-1, 1], H and W multiples of 16

latents = model.encode(images)  # [B, 96, H/16, W/16], whitened
semantic = model.semantic_channels(latents)  # [B, 32, H/16, W/16]
recon = model.decode(latents, images.shape[-2], images.shape[-1])
recon = recon.clamp(-1, 1)  # only for display or saving

from_pretrained(path_or_repo, *, device, dtype=torch.bfloat16, compile_encoder=True, compile_decoder=True, revision=None):

  • A Path is always a local directory.
  • A str is a local directory if one exists.
  • A str that looks like a path (it starts with ., / or ~, or has more than one /) but names no directory raises an error.
  • Any other str is a Hub repository id; only config.json and model.safetensors are downloaded.
  • The DINOv3-B weights ship inside model.safetensors, so a local copy loads offline.
  • Inference is CUDA only.

Details

  • Network: 164M parameters, 79M of them trained.
    • Encoder (87M): frozen DINOv3 ViT-B/16 behind a trainable patch embedding; all 12 blocks standardized, concatenated and projected linearly to 96 channels at stride 16.
    • Decoder (78M): 1×1 projection to width 1152, 4 transformer blocks, then a convolutional up-path from stride 16 to full resolution (512-channel handoff, levels of 256, 256, 128 and 128 channels, three residual 3 × 3 blocks per level).
  • Training data: about 14 million images from a mix of public and licensed or curated datasets: mostly photographs, plus book covers, a few text-heavy datasets and 1% synthetic rendered text. Aspect-ratio buckets, downsampled only. No training images are redistributed with the model.
  • Training:
    • Losses: mainly a DINOv3-B all-block feature MSE on the reconstruction, plus pixel and blurred-image MSE at about a tenth of its gradient; while the encoder trained, also a semantic alignment loss and VISReg.
    • Steps: about 150k at 256/384 px (batch 128) plus about 25k fine-tuning at five resolutions from 256 to 1024 px (batch 32). The free encoder weights (patch embedding and output projection) were trained jointly with the decoder for the first 52k steps and then frozen.
    • Matrix weights trained with Dion (dionw), the rest with AdamW. The released weights are EMA weights.
  • More: the technical report covers the semantic target, the decoder heads tried, loss weights, training phases, ablations and the benchmark protocol.
  • Related: semantic_vae, DINAC-AE-D2.

License

  • Code (the dinac3 package and scripts): Copyright 2026 data-archetype, licensed under the Apache License 2.0, see LICENSE.
  • Weights: the DINOv3 License, see LICENSE-DINOV3.md. The weights contain DINOv3-B's weights and a trained derivative of its patch embedding, so under the DINOv3 License (section 1.b.i) they may only be distributed under that licence, which must accompany them. Its terms include:
    • acknowledging the use of DINO Materials in publications (section 1.b.ii);
    • no reverse engineering (section 1.b.iv);
    • compliance with trade controls, which excludes ITAR-regulated uses and military, nuclear, espionage and weapons applications (sections 1.b.iii and 1.b.v).

The license: other tag above refers to the weights' DINOv3 License. See ATTRIBUTION.md for sources and design credits.

Citation

@misc{dinac3_96,
  title   = {dinac3_96: a DINOv3-based semantic autoencoder with a one-pass DINO-loss decoder},
  author  = {data-archetype},
  email   = {data-archetype@proton.me},
  year    = {2026},
  month   = oct,
  url     = {https://huggingface.co/data-archetype/dinac3_96},
}

[^fid]: Papers usually report FID on 50,000 square samples at a single resolution, so our numbers are not directly comparable with published ones.

[^rfid]: Papers usually report rFID on the 50,000 ImageNet validation images at a single square resolution, so our numbers are not directly comparable with published ones.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for data-archetype/dinac3_96

Finetuned
(5)
this model

Papers for data-archetype/dinac3_96