|
Download README.md from data-archetype/dinac3_96: direct link, hf CLI and curl.
- Browser
- Download file 10.9 kB
-
https://huggingface.co/data-archetype/dinac3_96/resolve/main/README.md
- Command line
-
hf download hf://data-archetype/dinac3_96/README.md
-
curl -L -o README.md https://huggingface.co/data-archetype/dinac3_96/resolve/main/README.md
10.9 kB
| license: other | |
| license_name: dinov3-license | |
| license_link: https://huggingface.co/data-archetype/dinac3_96/blob/main/LICENSE-DINOV3.md | |
| base_model: timm/vit_base_patch16_dinov3.lvd_1689m | |
| tags: | |
| - autoencoder | |
| - image-reconstruction | |
| - latent-space | |
| - latent-diffusion | |
| - dinov3 | |
| - pytorch | |
| # data-archetype/dinac3_96 | |
| **dinac3_96** is a deterministic semantic autoencoder for latent diffusion. It | |
| maps an RGB image to a 96-channel latent grid at stride 16 and decodes it back in | |
| a single forward pass. | |
| It is strongly inspired by [semantic_vae][semantic-vae], the latent with the | |
| fastest downstream DiT convergence we know of. dinac3_96 aims to keep that | |
| convergence while improving reconstruction and simplifying the encoder to a | |
| single DINOv3 pass instead of two. | |
| - **Encoder:** a frozen [DINOv3][dinov3] ViT-B/16 behind a trainable patch | |
| embedding. All twelve blocks are read, standardized per channel and mapped to | |
| the latent by one trainable linear projection. | |
| - **Latent:** 32 semantic channels, held during training to a fixed random | |
| projection of DINOv3-B's blocks summed with fitted block weights, plus 64 reconstruction channels (reconstruction | |
| losses and VISReg). No posterior noise. | |
| - **Decoder:** a 4-block transformer trunk at width 1152 followed by a dense | |
| convolutional up-path, a slightly modernized VQGAN-like decoder (stride 16 to | |
| full resolution, three residual blocks per level). | |
| - **Training losses:** mainly a DINOv3-B feature loss on the reconstruction, with | |
| pixel and blurred-image MSE terms each at about a tenth of its gradient, | |
| following semantic_vae's DINO-loss-trained decoder. | |
| **[Technical report](TECHNICAL_REPORT.md)** · | |
| **[Reconstruction gallery](https://huggingface.co/spaces/data-archetype/dinac3-results)** | |
| (39 images: original, reconstruction, RGB difference, latent PCA, per-image PSNR) | |
| ## Generation benchmark | |
| We evaluate a latent by training a class-conditional DiT on it (about 190M | |
| parameters, 100k steps at batch 256 on ImageNet-1k, flow matching with SPRINT | |
| routing). We score 10,000 samples against 810,000 real images on our | |
| mixed-resolution ImageNet benchmark: four aspect ratios at 256- and 384-px areas. | |
| Lower is better.[^fid] For each model: 50 NFEs, PDG 2.5–4.0, best settings kept. | |
| Bold marks the better value in each column. The DiT was trained on dinac3_96's latent (the released latent | |
| space, frozen since the joint phase), and its samples were decoded by the | |
| released decoder. | |
| | model | FID | MIND | Monge-DINO | | |
| |---|---:|---:|---:| | |
| | dinac3_96 | **`9.04`** | `8.29` | **`10.92`** | | |
| | semantic_vae | `10.02` | **`7.42`** | `15.94` | | |
| Monge-DINO uses DINOv3-B features, the same network whose features the decoder is | |
| trained to match, so it is not independent of the training loss; FID and MIND, on | |
| Inception features, are. The same holds for rMonge-DINOv3-B below. | |
| Our blind pairwise comparison of the two models' generations put them on equal | |
| footing. | |
| Protocol, DiT settings and caveats: | |
| [technical report, section 9.1](TECHNICAL_REPORT.md#91-generation-benchmark). | |
| ## Reconstruction | |
| **10,000 ImageNet images:** 10 real images per class with the generation | |
| benchmark's aspect-ratio quotas (9 at 256-px area and 1 at 384 px per class). | |
| Metrics compare the reconstructions with the same 10,000 originals.[^rfid] semantic_vae runs at its native bf16. | |
| | model | rFID | rMIND | rMonge-DINOv3-B | PSNR mean | SSIM | LPIPS-VGG | LPIPS-Alex | | |
| |---|---:|---:|---:|---:|---:|---:|---:| | |
| | dinac3_96 | **`0.809`** | **`0.166`** | **`1.205`** | **`28.16`** | **`0.8195`** | **`0.0970`** | **`0.0362`** | | |
| | semantic_vae | `2.149` | `0.820` | `5.994` | `25.19` | `0.7103` | `0.1715` | `0.0669` | | |
| PSNR here: reconstructions clamped and rounded to uint8, peak 255, mean of | |
| per-image values. | |
| **2k PSNR benchmark** (the image set of our earlier releases): | |
| | Model | Mean PSNR (dB) | Std (dB) | Median (dB) | P5 (dB) | P95 (dB) | | |
| |---|---:|---:|---:|---:|---:| | |
| | dinac3_96 | `32.15` | `5.11` | `31.64` | `24.27` | `40.70` | | |
| | semantic_vae | `28.03` | `4.58` | `27.79` | `20.98` | `35.44` | | |
| | dinac_ae_d2 | `35.59` | `4.87` | `35.40` | `27.89` | `43.51` | | |
| | FLUX.2 VAE | `36.28` | `4.53` | `36.07` | `28.89` | `43.63` | | |
| ## Latent interface | |
| - 96 channels at stride 16; image height and width must be multiples of 16. | |
| - Channels 0–63 are reconstruction channels, 64–95 semantic channels | |
| (`free_channels(z)`, `semantic_channels(z)`). | |
| - `encode(images)`: RGB in [-1, 1] → deterministic FP32 latents, whitened per | |
| channel with the shipped statistics (what our DiTs were trained on). | |
| - `decode(latents, height, width)`: whitened latents → FP32 RGB in about | |
| [-1, 1], unclamped. Any multiple of 16 works, including sizes above 1024 px. | |
| - `encode_raw` / `decode_raw` use the unwhitened latent; `whiten` / `dewhiten` | |
| convert. | |
| ## Precision | |
| - Set the dtype only through `from_pretrained(dtype=...)`; dtype casts after | |
| loading raise a `TypeError` (`.to(device)` works). | |
| - `torch.bfloat16` (default) runs under bf16 autocast, bit-identical to the | |
| training checkpoint under bf16 autocast. `torch.float32` runs without autocast. | |
| - Encoder and decoder are compiled by default (about 20 s per module on the | |
| first call). Pass `compile_encoder=False, compile_decoder=False` for eager. | |
| ## Speed | |
| RTX 5090, bf16, compiled (the default); mean of 20 timed batches after 3 | |
| warm-up batches. Peak VRAM includes the weights. | |
| | Resolution | Batch | Encode ms/image | Decode ms/image | Peak VRAM encode / decode (GiB) | | |
| |---:|---:|---:|---:|---:| | |
| | `256x256` | `128` | `0.41` | `1.51` | `3.3` / `8.5` | | |
| | `512x512` | `32` | `1.79` | `6.10` | `3.3` / `8.5` | | |
| | `1024x1024` | `8` | `9.03` | `25.4` | `3.3` / `8.5` | | |
| | `2048x2048` | `1` | `69.4` | `157` | `1.8` / `6.4` | | |
| Eager mode decodes about 2.4× slower (1.8× at 2048²) and peaks at 12.7 GiB on | |
| the batched rows. | |
| **Size limit.** No resolution limit is coded; GPU memory is the practical one. On | |
| a 32 GB RTX 5090, compiled decoding handles a single 4096 × 4096 image (15.8 s, | |
| 24.5 GiB peak); eager decoding runs out of memory at that size. | |
| ## Usage | |
| Download the repository and install the package, in an environment with a | |
| CUDA-enabled PyTorch (dependencies: `torch>=2.13`, `timm==1.0.26`, `safetensors`, | |
| `huggingface-hub`): | |
| ```bash | |
| hf download data-archetype/dinac3_96 --local-dir dinac3_96 | |
| pip install ./dinac3_96 | |
| ``` | |
| ```python | |
| import torch | |
| from dinac3 import Dinac3 | |
| model = Dinac3.from_pretrained( | |
| "data-archetype/dinac3_96", # or a local directory | |
| device="cuda", | |
| dtype=torch.bfloat16, | |
| ) | |
| images = ... # [B, 3, H, W] on CUDA, RGB in [-1, 1], H and W multiples of 16 | |
| latents = model.encode(images) # [B, 96, H/16, W/16], whitened | |
| semantic = model.semantic_channels(latents) # [B, 32, H/16, W/16] | |
| recon = model.decode(latents, images.shape[-2], images.shape[-1]) | |
| recon = recon.clamp(-1, 1) # only for display or saving | |
| ``` | |
| `from_pretrained(path_or_repo, *, device, dtype=torch.bfloat16, | |
| compile_encoder=True, compile_decoder=True, revision=None)`: | |
| - A `Path` is always a local directory. | |
| - A `str` is a local directory if one exists. | |
| - A `str` that looks like a path (it starts with `.`, `/` or `~`, or has more | |
| than one `/`) but names no directory raises an error. | |
| - Any other `str` is a Hub repository id; only `config.json` and | |
| `model.safetensors` are downloaded. | |
| - The DINOv3-B weights ship inside `model.safetensors`, so a local copy loads | |
| offline. | |
| - Inference is CUDA only. | |
| ## Details | |
| - **Network:** 164M parameters, 79M of them trained. | |
| - Encoder (87M): frozen DINOv3 ViT-B/16 behind a trainable patch embedding; | |
| all 12 blocks standardized, concatenated and projected linearly to 96 | |
| channels at stride 16. | |
| - Decoder (78M): 1×1 projection to width 1152, 4 transformer blocks, then a | |
| convolutional up-path from stride 16 to full resolution (512-channel | |
| handoff, levels of 256, 256, 128 and 128 channels, three residual 3 × 3 | |
| blocks per level). | |
| - **Training data:** about 14 million images from a mix of public and licensed | |
| or curated datasets: mostly photographs, plus book covers, a few text-heavy | |
| datasets and 1% synthetic rendered text. Aspect-ratio buckets, downsampled | |
| only. No training images are redistributed with the model. | |
| - **Training:** | |
| - Losses: mainly a DINOv3-B all-block feature MSE on the reconstruction, plus | |
| pixel and blurred-image MSE at about a tenth of its gradient; while the | |
| encoder trained, also a semantic alignment loss and [VISReg][visreg]. | |
| - Steps: about 150k at 256/384 px (batch 128) plus about 25k fine-tuning at five | |
| resolutions from 256 to 1024 px (batch 32). The free encoder weights (patch | |
| embedding and output projection) were trained jointly with the decoder for | |
| the first 52k steps and then frozen. | |
| - Matrix weights trained with Dion ([dionw][dionw]), the rest with AdamW. The | |
| released weights are EMA weights. | |
| - **More:** the [technical report](TECHNICAL_REPORT.md) covers the semantic | |
| target, the decoder heads tried, loss weights, training phases, ablations and | |
| the benchmark protocol. | |
| - **Related:** [semantic_vae][semantic-vae], | |
| [DINAC-AE-D2](https://huggingface.co/data-archetype/dinac_ae_d2). | |
| ## License | |
| - **Code** (the `dinac3` package and scripts): Copyright 2026 data-archetype, | |
| licensed under the Apache License 2.0, see [LICENSE](LICENSE). | |
| - **Weights:** the DINOv3 License, see [LICENSE-DINOV3.md](LICENSE-DINOV3.md). | |
| The weights contain DINOv3-B's weights and a trained derivative of its patch | |
| embedding, so under the DINOv3 License (section 1.b.i) they may only be | |
| distributed under that licence, which must accompany them. Its terms include: | |
| - acknowledging the use of DINO Materials in publications (section 1.b.ii); | |
| - no reverse engineering (section 1.b.iv); | |
| - compliance with trade controls, which excludes ITAR-regulated uses and | |
| military, nuclear, espionage and weapons applications (sections 1.b.iii and | |
| 1.b.v). | |
| The `license: other` tag above refers to the weights' DINOv3 License. See | |
| [ATTRIBUTION.md](ATTRIBUTION.md) for sources and design credits. | |
| ## Citation | |
| ```bibtex | |
| @misc{dinac3_96, | |
| title = {dinac3_96: a DINOv3-based semantic autoencoder with a one-pass DINO-loss decoder}, | |
| author = {data-archetype}, | |
| email = {data-archetype@proton.me}, | |
| year = {2026}, | |
| month = oct, | |
| url = {https://huggingface.co/data-archetype/dinac3_96}, | |
| } | |
| ``` | |
| [dinov3]: https://arxiv.org/abs/2508.10104 | |
| [^fid]: Papers usually report FID on 50,000 square samples at a single | |
| resolution, so our numbers are not directly comparable with published ones. | |
| [^rfid]: Papers usually report rFID on the 50,000 ImageNet validation images at | |
| a single square resolution, so our numbers are not directly comparable with | |
| published ones. | |
| [semantic-vae]: https://huggingface.co/well9472/semantic_vae | |
| [visreg]: https://arxiv.org/abs/2606.02572 | |
| [dionw]: https://github.com/JTriggerFish/dionw | |