qat-transfer checkpoints: timm supervised ViT / DeiT-III / Swin

Finetuned checkpoints for Zero-Shot Quantization via Weight-Space Arithmetic (NeurIPS 2026). Code: github.com/gladia-research-group/qat-transfer · arXiv · OpenReview · Research graph

For every backbone and dataset this repo has a pair of checkpoints that share a key set exactly:

  • FP: full-precision finetune.
  • QAT: 3-bit, per-channel quantization-aware finetune (STE), saved with the QAT wrappers stripped.

Their difference is the quantization vector QV = QAT - FP. Adding a donor task's QV to a different receiver task's FP checkpoint, FP_receiver + lambda * QV_donor, recovers much of the receiver's QAT benefit under 3-bit post-training quantization, with no receiver-side QAT and, at lambda = 1, no receiver data.

The other model families are in gladia/qat-transfer-timm, gladia/qat-transfer-open_clip and gladia/qat-transfer-text.

Contents

Backbones: deit3_base_patch16_224.fb_in1k, deit3_large_patch16_224.fb_in1k, swin_base_patch4_window7_224.ms_in22k_ft_in1k, swin_large_patch4_window7_224.ms_in22k_ft_in1k, vit_base_patch16_224.orig_in21k, vit_large_patch16_224.orig_in21k, vit_huge_patch14_224.orig_in21k.

Datasets: Cars, DTD, EuroSAT, GTSRB, MNIST, RESISC45, SUN397, SVHN, CIFAR10, CIFAR100, STL10, Food101, Flowers102, FER2013, PCAM, OxfordIIITPet, RenderedSST2, EMNIST, FashionMNIST, KMNIST, TinyImageNet, ImageNet.

Shared configuration: AdamW, lr=1e-5, wd=0.1, no label smoothing, gradient clipping at 1.0, seed 2038, the 1x training schedule (mult=1). QAT and evaluation-time PTQ both use 3 bits, per-channel, with the classification head (head) left unquantized. Larger models were trained at a smaller per-device batch with gradient accumulation; the optim= directory names the per-device batch size.

Each run directory holds classifier_epoch_{N}.pt (full classifier state dict, backbone + head) and head_epoch_{N}.pt (the head, saved as a pickled torch.nn.Linear; load with weights_only=False). N is the dataset's reference epoch count, and run_meta.json, where present, records the realized schedule.

PV-Tuning donors

checkpoints/vision/ilharco_timm_supervised/pv/vit_base_patch16_224_orig_in21k/ holds PV-Tuning finetunes (delta=0.0, tau=0.01) used as alternative donors in the PV-Tuning transfer experiment (008_pv_transfer). Alongside the settled classifier, each run has a pv_state_epoch_{N}.pt sidecar with the per-layer {codes, scale, latent}.

Only 5 of the 22 PV-Tuning donors are included (Cars, EuroSAT, Flowers102, RESISC45, STL10). The other 17 were trained and used for the reported results, but their checkpoint files were lost to a storage failure afterwards. They can be regenerated exactly as the originals were produced with code/src/vision/ilharco_timm_supervised/finetune_pv.py in the code repository (same config, seed 2038); the FP and QAT checkpoints in this repository are complete.

Layout

The repo root mirrors the code's storage/ directory, so the scripts in the code repository read these files with no changes:

checkpoints/vision/ilharco_timm_supervised/fp/{model}/{dataset}/optim=.../mult=1/seed=2038/
checkpoints/vision/ilharco_timm_supervised/qat/{model}/{dataset}/optim=.../mult=1/qat=bits=3_gran=channel_skip=.../seed=2038/

Download

pip install -U huggingface_hub
hf download gladia/qat-transfer-timm --local-dir storage

Download one backbone only:

hf download gladia/qat-transfer-timm --local-dir storage --include "*/vit_base_patch16_224_orig_in21k/*"

Then set CHECKPOINT_BASE_PATH=storage/checkpoints and HEAD_BASE_PATH=storage/heads in the code repository's .env. All three repos can be downloaded into the same storage/.

Building a quantization vector by hand

The checkpoints are plain PyTorch state dicts, so the method needs nothing beyond torch:

import glob, torch

def load(pattern):
    (path,) = glob.glob(pattern)
    return torch.load(path, map_location="cpu")

root = "storage/checkpoints/vision/ilharco_timm_supervised"
fp_donor  = load(f"{root}/fp/vit_base_patch16_224_orig_in21k/EuroSAT/*/mult=1/seed=2038/classifier_epoch_*.pt")
qat_donor = load(f"{root}/qat/vit_base_patch16_224_orig_in21k/EuroSAT/*/mult=1/qat=*/seed=2038/classifier_epoch_*.pt")
fp_recv   = load(f"{root}/fp/vit_base_patch16_224_orig_in21k/DTD/*/mult=1/seed=2038/classifier_epoch_*.pt")

lam = 1.0
# the head is task-specific: keep the receiver's own
patched = {k: fp_recv[k] if k.startswith("model.head") else fp_recv[k] + lam * (qat_donor[k] - fp_donor[k]) for k in fp_recv}

Load patched into the receiver model, apply 3-bit per-channel PTQ to every linear layer except the head, and evaluate. code/experiments/*/001_qat_transfer/qv_transfer.py in the code repository does exactly this.

Research graph

The questions behind each experiment phase, how they depend on one another and on the paper's propositions, and the scripts that answer them are laid out as a public Flywheel graph. Each node names its question, its method and the code in the code repository that implements it.

License

Finetuned from timm checkpoints released under Apache-2.0 (ViT, DeiT-III) and MIT (Swin). Check each upstream model card before redistribution. The finetuning datasets carry their own terms.

Citation

@inproceedings{solombrino2026zeroshot,
  title     = {Zero-Shot Quantization via Weight-Space Arithmetic},
  author    = {Solombrino, Daniele and Gargiulo, Antonio Andrea and Zirilli, Alessandro and Zhou, Luca and Minut, Adrian Robert and Rodol{\`a}, Emanuele},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026},
  url       = {https://openreview.net/forum?id=wrUEnnSgXa}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for gladia/qat-transfer-timm