dinac3_96 / README.md
data-archetype's picture
Card: drop a redundant sentence
9028103 verified
|
Raw History Blame Contribute Delete
10.9 kB
---
license: other
license_name: dinov3-license
license_link: https://huggingface.co/data-archetype/dinac3_96/blob/main/LICENSE-DINOV3.md
base_model: timm/vit_base_patch16_dinov3.lvd_1689m
tags:
- autoencoder
- image-reconstruction
- latent-space
- latent-diffusion
- dinov3
- pytorch
---
# data-archetype/dinac3_96
**dinac3_96** is a deterministic semantic autoencoder for latent diffusion. It
maps an RGB image to a 96-channel latent grid at stride 16 and decodes it back in
a single forward pass.
It is strongly inspired by [semantic_vae][semantic-vae], the latent with the
fastest downstream DiT convergence we know of. dinac3_96 aims to keep that
convergence while improving reconstruction and simplifying the encoder to a
single DINOv3 pass instead of two.
- **Encoder:** a frozen [DINOv3][dinov3] ViT-B/16 behind a trainable patch
embedding. All twelve blocks are read, standardized per channel and mapped to
the latent by one trainable linear projection.
- **Latent:** 32 semantic channels, held during training to a fixed random
projection of DINOv3-B's blocks summed with fitted block weights, plus 64 reconstruction channels (reconstruction
losses and VISReg). No posterior noise.
- **Decoder:** a 4-block transformer trunk at width 1152 followed by a dense
convolutional up-path, a slightly modernized VQGAN-like decoder (stride 16 to
full resolution, three residual blocks per level).
- **Training losses:** mainly a DINOv3-B feature loss on the reconstruction, with
pixel and blurred-image MSE terms each at about a tenth of its gradient,
following semantic_vae's DINO-loss-trained decoder.
**[Technical report](TECHNICAL_REPORT.md)** ·
**[Reconstruction gallery](https://huggingface.co/spaces/data-archetype/dinac3-results)**
(39 images: original, reconstruction, RGB difference, latent PCA, per-image PSNR)
## Generation benchmark
We evaluate a latent by training a class-conditional DiT on it (about 190M
parameters, 100k steps at batch 256 on ImageNet-1k, flow matching with SPRINT
routing). We score 10,000 samples against 810,000 real images on our
mixed-resolution ImageNet benchmark: four aspect ratios at 256- and 384-px areas.
Lower is better.[^fid] For each model: 50 NFEs, PDG 2.5–4.0, best settings kept.
Bold marks the better value in each column. The DiT was trained on dinac3_96's latent (the released latent
space, frozen since the joint phase), and its samples were decoded by the
released decoder.
| model | FID | MIND | Monge-DINO |
|---|---:|---:|---:|
| dinac3_96 | **`9.04`** | `8.29` | **`10.92`** |
| semantic_vae | `10.02` | **`7.42`** | `15.94` |
Monge-DINO uses DINOv3-B features, the same network whose features the decoder is
trained to match, so it is not independent of the training loss; FID and MIND, on
Inception features, are. The same holds for rMonge-DINOv3-B below.
Our blind pairwise comparison of the two models' generations put them on equal
footing.
Protocol, DiT settings and caveats:
[technical report, section 9.1](TECHNICAL_REPORT.md#91-generation-benchmark).
## Reconstruction
**10,000 ImageNet images:** 10 real images per class with the generation
benchmark's aspect-ratio quotas (9 at 256-px area and 1 at 384 px per class).
Metrics compare the reconstructions with the same 10,000 originals.[^rfid] semantic_vae runs at its native bf16.
| model | rFID | rMIND | rMonge-DINOv3-B | PSNR mean | SSIM | LPIPS-VGG | LPIPS-Alex |
|---|---:|---:|---:|---:|---:|---:|---:|
| dinac3_96 | **`0.809`** | **`0.166`** | **`1.205`** | **`28.16`** | **`0.8195`** | **`0.0970`** | **`0.0362`** |
| semantic_vae | `2.149` | `0.820` | `5.994` | `25.19` | `0.7103` | `0.1715` | `0.0669` |
PSNR here: reconstructions clamped and rounded to uint8, peak 255, mean of
per-image values.
**2k PSNR benchmark** (the image set of our earlier releases):
| Model | Mean PSNR (dB) | Std (dB) | Median (dB) | P5 (dB) | P95 (dB) |
|---|---:|---:|---:|---:|---:|
| dinac3_96 | `32.15` | `5.11` | `31.64` | `24.27` | `40.70` |
| semantic_vae | `28.03` | `4.58` | `27.79` | `20.98` | `35.44` |
| dinac_ae_d2 | `35.59` | `4.87` | `35.40` | `27.89` | `43.51` |
| FLUX.2 VAE | `36.28` | `4.53` | `36.07` | `28.89` | `43.63` |
## Latent interface
- 96 channels at stride 16; image height and width must be multiples of 16.
- Channels 0–63 are reconstruction channels, 64–95 semantic channels
(`free_channels(z)`, `semantic_channels(z)`).
- `encode(images)`: RGB in [-1, 1] → deterministic FP32 latents, whitened per
channel with the shipped statistics (what our DiTs were trained on).
- `decode(latents, height, width)`: whitened latents → FP32 RGB in about
[-1, 1], unclamped. Any multiple of 16 works, including sizes above 1024 px.
- `encode_raw` / `decode_raw` use the unwhitened latent; `whiten` / `dewhiten`
convert.
## Precision
- Set the dtype only through `from_pretrained(dtype=...)`; dtype casts after
loading raise a `TypeError` (`.to(device)` works).
- `torch.bfloat16` (default) runs under bf16 autocast, bit-identical to the
training checkpoint under bf16 autocast. `torch.float32` runs without autocast.
- Encoder and decoder are compiled by default (about 20 s per module on the
first call). Pass `compile_encoder=False, compile_decoder=False` for eager.
## Speed
RTX 5090, bf16, compiled (the default); mean of 20 timed batches after 3
warm-up batches. Peak VRAM includes the weights.
| Resolution | Batch | Encode ms/image | Decode ms/image | Peak VRAM encode / decode (GiB) |
|---:|---:|---:|---:|---:|
| `256x256` | `128` | `0.41` | `1.51` | `3.3` / `8.5` |
| `512x512` | `32` | `1.79` | `6.10` | `3.3` / `8.5` |
| `1024x1024` | `8` | `9.03` | `25.4` | `3.3` / `8.5` |
| `2048x2048` | `1` | `69.4` | `157` | `1.8` / `6.4` |
Eager mode decodes about 2.4× slower (1.8× at 2048²) and peaks at 12.7 GiB on
the batched rows.
**Size limit.** No resolution limit is coded; GPU memory is the practical one. On
a 32 GB RTX 5090, compiled decoding handles a single 4096 × 4096 image (15.8 s,
24.5 GiB peak); eager decoding runs out of memory at that size.
## Usage
Download the repository and install the package, in an environment with a
CUDA-enabled PyTorch (dependencies: `torch>=2.13`, `timm==1.0.26`, `safetensors`,
`huggingface-hub`):
```bash
hf download data-archetype/dinac3_96 --local-dir dinac3_96
pip install ./dinac3_96
```
```python
import torch
from dinac3 import Dinac3
model = Dinac3.from_pretrained(
"data-archetype/dinac3_96", # or a local directory
device="cuda",
dtype=torch.bfloat16,
)
images = ... # [B, 3, H, W] on CUDA, RGB in [-1, 1], H and W multiples of 16
latents = model.encode(images) # [B, 96, H/16, W/16], whitened
semantic = model.semantic_channels(latents) # [B, 32, H/16, W/16]
recon = model.decode(latents, images.shape[-2], images.shape[-1])
recon = recon.clamp(-1, 1) # only for display or saving
```
`from_pretrained(path_or_repo, *, device, dtype=torch.bfloat16,
compile_encoder=True, compile_decoder=True, revision=None)`:
- A `Path` is always a local directory.
- A `str` is a local directory if one exists.
- A `str` that looks like a path (it starts with `.`, `/` or `~`, or has more
than one `/`) but names no directory raises an error.
- Any other `str` is a Hub repository id; only `config.json` and
`model.safetensors` are downloaded.
- The DINOv3-B weights ship inside `model.safetensors`, so a local copy loads
offline.
- Inference is CUDA only.
## Details
- **Network:** 164M parameters, 79M of them trained.
- Encoder (87M): frozen DINOv3 ViT-B/16 behind a trainable patch embedding;
all 12 blocks standardized, concatenated and projected linearly to 96
channels at stride 16.
- Decoder (78M): 1×1 projection to width 1152, 4 transformer blocks, then a
convolutional up-path from stride 16 to full resolution (512-channel
handoff, levels of 256, 256, 128 and 128 channels, three residual 3 × 3
blocks per level).
- **Training data:** about 14 million images from a mix of public and licensed
or curated datasets: mostly photographs, plus book covers, a few text-heavy
datasets and 1% synthetic rendered text. Aspect-ratio buckets, downsampled
only. No training images are redistributed with the model.
- **Training:**
- Losses: mainly a DINOv3-B all-block feature MSE on the reconstruction, plus
pixel and blurred-image MSE at about a tenth of its gradient; while the
encoder trained, also a semantic alignment loss and [VISReg][visreg].
- Steps: about 150k at 256/384 px (batch 128) plus about 25k fine-tuning at five
resolutions from 256 to 1024 px (batch 32). The free encoder weights (patch
embedding and output projection) were trained jointly with the decoder for
the first 52k steps and then frozen.
- Matrix weights trained with Dion ([dionw][dionw]), the rest with AdamW. The
released weights are EMA weights.
- **More:** the [technical report](TECHNICAL_REPORT.md) covers the semantic
target, the decoder heads tried, loss weights, training phases, ablations and
the benchmark protocol.
- **Related:** [semantic_vae][semantic-vae],
[DINAC-AE-D2](https://huggingface.co/data-archetype/dinac_ae_d2).
## License
- **Code** (the `dinac3` package and scripts): Copyright 2026 data-archetype,
licensed under the Apache License 2.0, see [LICENSE](LICENSE).
- **Weights:** the DINOv3 License, see [LICENSE-DINOV3.md](LICENSE-DINOV3.md).
The weights contain DINOv3-B's weights and a trained derivative of its patch
embedding, so under the DINOv3 License (section 1.b.i) they may only be
distributed under that licence, which must accompany them. Its terms include:
- acknowledging the use of DINO Materials in publications (section 1.b.ii);
- no reverse engineering (section 1.b.iv);
- compliance with trade controls, which excludes ITAR-regulated uses and
military, nuclear, espionage and weapons applications (sections 1.b.iii and
1.b.v).
The `license: other` tag above refers to the weights' DINOv3 License. See
[ATTRIBUTION.md](ATTRIBUTION.md) for sources and design credits.
## Citation
```bibtex
@misc{dinac3_96,
title = {dinac3_96: a DINOv3-based semantic autoencoder with a one-pass DINO-loss decoder},
author = {data-archetype},
email = {data-archetype@proton.me},
year = {2026},
month = oct,
url = {https://huggingface.co/data-archetype/dinac3_96},
}
```
[dinov3]: https://arxiv.org/abs/2508.10104
[^fid]: Papers usually report FID on 50,000 square samples at a single
resolution, so our numbers are not directly comparable with published ones.
[^rfid]: Papers usually report rFID on the 50,000 ImageNet validation images at
a single square resolution, so our numbers are not directly comparable with
published ones.
[semantic-vae]: https://huggingface.co/well9472/semantic_vae
[visreg]: https://arxiv.org/abs/2606.02572
[dionw]: https://github.com/JTriggerFish/dionw