qat-transfer checkpoints: OpenCLIP (LAION-2B) ViT-B/16, ViT-L/14, ViT-H/14

Finetuned checkpoints for Zero-Shot Quantization via Weight-Space Arithmetic (NeurIPS 2026). Code: github.com/gladia-research-group/qat-transfer · arXiv · OpenReview · Research graph

For every backbone and dataset this repo has a pair of checkpoints that share a key set exactly:

  • FP: full-precision finetune.
  • QAT: 3-bit, per-channel quantization-aware finetune (STE), saved with the QAT wrappers stripped.

Their difference is the quantization vector QV = QAT - FP. Adding a donor task's QV to a different receiver task's FP checkpoint, FP_receiver + lambda * QV_donor, recovers much of the receiver's QAT benefit under 3-bit post-training quantization, with no receiver-side QAT and, at lambda = 1, no receiver data.

The other model families are in gladia/qat-transfer-timm, gladia/qat-transfer-open_clip and gladia/qat-transfer-text.

Contents

Backbones: ViT-B-16 / laion2b_s34b_b88k, ViT-L-14 / laion2b_s32b_b82k, ViT-H-14 / laion2b_s32b_b79k.

Datasets: Cars, DTD, EuroSAT, GTSRB, MNIST, RESISC45, SUN397, SVHN, CIFAR10, CIFAR100, STL10, Food101, Flowers102, FER2013, PCAM, OxfordIIITPet, RenderedSST2, EMNIST, FashionMNIST, KMNIST, TinyImageNet, ImageNet.

Shared configuration: AdamW, lr=1e-5, wd=0.1, no label smoothing, gradient clipping at 1.0, seed 2038, the 1x training schedule (mult=1). QAT and evaluation-time PTQ both use 3 bits, per-channel, with the classification head (classification_head) left unquantized. Larger models were trained at a smaller per-device batch with gradient accumulation; the optim= directory names the per-device batch size.

Each run directory holds a single epoch_{N}.pt: the image encoder's state dict, without a head. The classification head is the frozen zero-shot head built from the text tower, stored once per model and dataset under heads/vision/ilharco_open_clip/{model}/head_{dataset}.pt (weight, bias, normalize_flag); point HEAD_BASE_PATH at storage/heads. N is the dataset's reference epoch count, and run_meta.json, where present, records the realized schedule.

Layout

The repo root mirrors the code's storage/ directory, so the scripts in the code repository read these files with no changes:

checkpoints/vision/ilharco_open_clip/fp/{model}/{dataset}/optim=.../mult=1/seed=2038/
checkpoints/vision/ilharco_open_clip/qat/{model}/{dataset}/optim=.../mult=1/qat=bits=3_gran=channel_skip=.../seed=2038/

Download

pip install -U huggingface_hub
hf download gladia/qat-transfer-open_clip --local-dir storage

Download one backbone only:

hf download gladia/qat-transfer-open_clip --local-dir storage --include "*/ViT_B_16__laion2b_s34b_b88k/*"

Then set CHECKPOINT_BASE_PATH=storage/checkpoints and HEAD_BASE_PATH=storage/heads in the code repository's .env. All three repos can be downloaded into the same storage/.

Building a quantization vector by hand

The checkpoints are plain PyTorch state dicts, so the method needs nothing beyond torch:

import glob, torch

def load(pattern):
    (path,) = glob.glob(pattern)
    return torch.load(path, map_location="cpu")

root = "storage/checkpoints/vision/ilharco_open_clip"
fp_donor  = load(f"{root}/fp/ViT_B_16__laion2b_s34b_b88k/EuroSAT/*/mult=1/seed=2038/epoch_*.pt")
qat_donor = load(f"{root}/qat/ViT_B_16__laion2b_s34b_b88k/EuroSAT/*/mult=1/qat=*/seed=2038/epoch_*.pt")
fp_recv   = load(f"{root}/fp/ViT_B_16__laion2b_s34b_b88k/DTD/*/mult=1/seed=2038/epoch_*.pt")

lam = 1.0
patched = {k: fp_recv[k] + lam * (qat_donor[k] - fp_donor[k]) for k in fp_recv}

Load patched into the receiver model, apply 3-bit per-channel PTQ to every linear layer except the head, and evaluate. code/experiments/*/001_qat_transfer/qv_transfer.py in the code repository does exactly this.

Research graph

The questions behind each experiment phase, how they depend on one another and on the paper's propositions, and the scripts that answer them are laid out as a public Flywheel graph. Each node names its question, its method and the code in the code repository that implements it.

License

Finetuned from OpenCLIP LAION-2B checkpoints (MIT). The finetuning datasets carry their own terms.

Citation

@inproceedings{solombrino2026zeroshot,
  title     = {Zero-Shot Quantization via Weight-Space Arithmetic},
  author    = {Solombrino, Daniele and Gargiulo, Antonio Andrea and Zirilli, Alessandro and Zhou, Luca and Minut, Adrian Robert and Rodol{\`a}, Emanuele},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026},
  url       = {https://openreview.net/forum?id=wrUEnnSgXa}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for gladia/qat-transfer-open_clip