PoolDINO
Pooling Representation Autoencoders for Efficient Diffusion
Ramón Calvo-González · Youssef Saied · François Fleuret
University of Geneva · Meta
Paper · Code · Project page
PoolDINO compresses pretrained visual representations with a learned spatial pooling operator. It merges each local window of DINOv3 tokens into one token, reducing the sequence length modeled by the diffusion Transformer. The pooling operator is trained jointly with the RGB decoder, preserving the standard two-stage representation-autoencoder training procedure.
This repository contains JAX/Orbax checkpoints for ImageNet-256 RGB reconstruction and class-conditional image generation. These are not Diffusers or Transformers pipeline checkpoints; use the PoolDINO implementation.
Selected examples. Learned 2 × 2 pooling, 180 training epochs, 100 Euler steps, internal guidance (IG) scale 2.00, no classifier-free guidance (CFG).
How it works
- Train the image decoder. A frozen DINOv3-L/16 encoder produces a 16 × 16 grid. A shared affine map pools each non-overlapping window into one 1,024-dimensional token. Tokens are repeated over their original windows before the ViT-XL RGB decoder reconstructs the image. The pooling operator and decoder are trained together.
- Train the generator. Freeze the encoder, pooling operator, and RGB decoder. Train a flow-matching Transformer directly on the normalized, compressed latent grid. Internal guidance uses an intermediate generator prediction to guide sampling.
Available checkpoints
All experiment directory names below end in -dinol-vitxl-raev2official-tfds. The historical identifier repeatconv means learned spatial pooling; pool1x1 is the unpooled reference.
| Experiment prefix | Pooling | Unique tokens | Token compression | Generator epochs | Generator step |
|---|---|---|---|---|---|
pool1x1 |
Unpooled | 256 | 1× | 80 | 100080 |
repeatconv2x2 |
Learned 2 × 2 | 64 | 4× | 80 | 100080 |
repeatconv2x2 |
Learned 2 × 2 | 64 | 4× | 180 | 225180 |
repeatconv2x4 |
Learned 2 × 4 | 32 | 8× | 80 | 100080 |
repeatconv4x2 |
Learned 4 × 2 | 32 | 8× | 80 | 100080 |
repeatconv4x4 |
Learned 4 × 4 | 16 | 16× | 80 | 100080 |
Generator checkpoints are under pooled-generator/<experiment>/<step>/. Their matching RGB decoder checkpoints are under pooled-decoder/<experiment>/40032/, with pooled_latent_stats.npz in the decoder experiment directory.
An additional learned 1 × 1 decoder (repeatconv1x1) is included. It is not the unpooled pool1x1 reference and has no corresponding generator in this release.
The 300-epoch 4 × 4 generator discussed in the paper is not included in the current upload. Neither are the average-pooling generators or the later 222-/390-epoch runs.
Generation results
Reported ImageNet-256 results with 100 Euler steps, 50,000 samples, EMA weights, and IG alone. The scales below reproduce the settings reported in the main results tables, rather than asserting the optimum over every subsequent fine-scale sweep.
| Model | Epochs | IG scale | FID ↓ | Inception Score ↑ |
|---|---|---|---|---|
| Unpooled reference | 80 | 1.75 | 1.08 | 262.00 |
| Learned 2 × 2 | 80 | 1.75 | 1.09 | 250.36 |
| Learned 2 × 4 | 80 | 2.00 | 1.19 | 249.58 |
| Learned 4 × 2 | 80 | 2.00 | 1.21 | 249.96 |
| Learned 4 × 4 | 80 | 2.75 | 1.44 | 247.07 |
| Learned 2 × 2 | 180 | 2.00 | 1.05 | 272.00 |
At 100 sampling steps, 4× and 16× token compression increase latent-sampling throughput by approximately 3.7× and 9.0× relative to the unpooled reference. These measurements use an NVIDIA H100 NVL at batch size 128 and exclude RGB decoding. The 180-epoch run is an extended-training experiment, not a FLOP-matched run.
Download
Install the Hugging Face CLI and authenticate if the repository requires access:
pip install -U huggingface_hub
hf auth login
Download one matching generator, decoder, and latent-statistics set. For example, the 80-epoch 2 × 2 model:
hf download noctrog/pooldino \
--include 'pooled-decoder/repeatconv2x2-dinol-vitxl-raev2official-tfds/*' \
--include 'pooled-generator/repeatconv2x2-dinol-vitxl-raev2official-tfds/100080/*' \
--local-dir checkpoints/pooldino
For the 180-epoch generator, replace 100080 with 225180; it uses the same decoder. Keep the complete checkpoint subtree, including metadata and sharded array files. Downloading a single manifest is not sufficient.
Sampling
The implementation is maintained in the code repository. The intended GPU environment is Linux with CUDA 12 and Python 3.13.
From the code repository, install the locked environment:
uv sync --locked --extra cuda --extra metrics
The following example assumes the download above is in checkpoints/pooldino relative to the code repository. Supply the ImageNet validation-label file described in the code README:
uv run python -m pooldino.eval.gfid_pooled_decoder_adm \
--generator-path checkpoints/pooldino/pooled-generator/repeatconv2x2-dinol-vitxl-raev2official-tfds \
--generator-step 100080 \
--pooled-decoder-path checkpoints/pooldino/pooled-decoder/repeatconv2x2-dinol-vitxl-raev2official-tfds \
--pooled-decoder-step 40032 \
--condition-labels-path /path/to/imagenet2012-validation-labels-tfds.npz \
--protocol raev2_ig --ig-scale 1.75 --generate-only
- Point
--generator-pathand--pooled-decoder-pathto the experiment directories, not the numeric step directories. Explicitly select the step if multiple checkpoints are downloaded. - Keep each decoder's matching
pooled_latent_stats.npz. If restoration reports a source-path mismatch after relocation, follow the code README'sPOOLDINO_ARTIFACT_PATH_MAPinstructions; do not disable artifact-identity checks. raev2_igselects the 100-step evaluation protocol, seed 42, 50,000 images, and BF16 inference.--generate-onlyskips metric computation, not generation of the full evaluation sample set.- IG is active on
[0, 0.9]in the implementation's noise-to-data time convention. IG and CFG are neutral at scale 1. The encoder-reconstruction guidance used in appendix ablations has a different scale convention. - Preserve the checkpoint's time-shift settings. Check the saved
generation_config.txtfor the effective sampling protocol. - Frozen encoder weights are resolved separately by the implementation. Obtain pretrained weights and datasets under their respective access conditions and terms.
The instructions above are based on the implementation and uploaded checkpoint layout; they are not a claim of a fresh end-to-end GPU reproduction of this Hub download.
Intended use and limitations
These checkpoints support research on representation autoencoders, spatial compression, and class-conditional ImageNet image generation at 256 × 256. They are not text-to-image models.
Compression trades spatial detail and transfer performance for faster generation. Comparable guided generation quality at 4× compression does not imply that all semantic information is preserved: learned pooling gives lower classification-probe accuracy than simpler baselines and does not consistently outperform averaging on segmentation and depth estimation. Stronger compression can produce visible fine-detail artifacts.
Generated images can reflect biases and limitations of ImageNet and the pretrained encoder. The models have not been validated for safety-critical or unrestricted deployment. The selected examples above should not substitute for distribution-level evaluation; the paper includes random sample grids without quality filtering.
Terms and attribution
This model card does not assign a new license to the checkpoints. Upstream components and datasets have their own terms; repository access should not be interpreted as unrestricted commercial-use permission.
PoolDINO builds on the representation-autoencoder framework and RAEv2. The paper and code provide the full methodology, experimental protocols, and upstream citations.
Citation
@misc{calvogonzález2026poolingrepresentationautoencodersefficient,
title={Pooling Representation Autoencoders for Efficient Diffusion},
author={Ramón Calvo-González and Youssef Saied and François Fleuret},
year={2026},
eprint={2610.09242},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.09242},
}