Instructions to use gladia/qat-transfer-open_clip with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- OpenCLIP
How to use gladia/qat-transfer-open_clip with OpenCLIP:
import open_clip model, preprocess_train, preprocess_val = open_clip.create_model_and_transforms('hf-hub:gladia/qat-transfer-open_clip') tokenizer = open_clip.get_tokenizer('hf-hub:gladia/qat-transfer-open_clip') - Notebooks
- Google Colab
- Kaggle
qat-transfer checkpoints: OpenCLIP (LAION-2B) ViT-B/16, ViT-L/14, ViT-H/14
Finetuned checkpoints for Zero-Shot Quantization via Weight-Space Arithmetic (NeurIPS 2026). Code: github.com/gladia-research-group/qat-transfer · arXiv · OpenReview · Research graph
For every backbone and dataset this repo has a pair of checkpoints that share a key set exactly:
- FP: full-precision finetune.
- QAT: 3-bit, per-channel quantization-aware finetune (STE), saved with the QAT wrappers stripped.
Their difference is the quantization vector QV = QAT - FP. Adding a donor task's QV to a different receiver task's FP checkpoint, FP_receiver + lambda * QV_donor, recovers much of the receiver's QAT benefit under 3-bit post-training quantization, with no receiver-side QAT and, at lambda = 1, no receiver data.
The other model families are in gladia/qat-transfer-timm, gladia/qat-transfer-open_clip and gladia/qat-transfer-text.
Contents
Backbones: ViT-B-16 / laion2b_s34b_b88k, ViT-L-14 / laion2b_s32b_b82k, ViT-H-14 / laion2b_s32b_b79k.
Datasets: Cars, DTD, EuroSAT, GTSRB, MNIST, RESISC45, SUN397, SVHN, CIFAR10, CIFAR100, STL10, Food101, Flowers102, FER2013, PCAM, OxfordIIITPet, RenderedSST2, EMNIST, FashionMNIST, KMNIST, TinyImageNet, ImageNet.
Shared configuration: AdamW, lr=1e-5, wd=0.1, no label smoothing, gradient clipping at 1.0, seed 2038, the 1x training schedule (mult=1). QAT and evaluation-time PTQ both use 3 bits, per-channel, with the classification head (classification_head) left unquantized. Larger models were trained at a smaller per-device batch with gradient accumulation; the optim= directory names the per-device batch size.
Each run directory holds a single epoch_{N}.pt: the image encoder's state dict, without a head. The classification head is the frozen zero-shot head built from the text tower, stored once per model and dataset under heads/vision/ilharco_open_clip/{model}/head_{dataset}.pt (weight, bias, normalize_flag); point HEAD_BASE_PATH at storage/heads. N is the dataset's reference epoch count, and run_meta.json, where present, records the realized schedule.
Layout
The repo root mirrors the code's storage/ directory, so the scripts in the code repository read these files with no changes:
checkpoints/vision/ilharco_open_clip/fp/{model}/{dataset}/optim=.../mult=1/seed=2038/
checkpoints/vision/ilharco_open_clip/qat/{model}/{dataset}/optim=.../mult=1/qat=bits=3_gran=channel_skip=.../seed=2038/
Download
pip install -U huggingface_hub
hf download gladia/qat-transfer-open_clip --local-dir storage
Download one backbone only:
hf download gladia/qat-transfer-open_clip --local-dir storage --include "*/ViT_B_16__laion2b_s34b_b88k/*"
Then set CHECKPOINT_BASE_PATH=storage/checkpoints and HEAD_BASE_PATH=storage/heads in the code repository's .env. All three repos can be downloaded into the same storage/.
Building a quantization vector by hand
The checkpoints are plain PyTorch state dicts, so the method needs nothing beyond torch:
import glob, torch
def load(pattern):
(path,) = glob.glob(pattern)
return torch.load(path, map_location="cpu")
root = "storage/checkpoints/vision/ilharco_open_clip"
fp_donor = load(f"{root}/fp/ViT_B_16__laion2b_s34b_b88k/EuroSAT/*/mult=1/seed=2038/epoch_*.pt")
qat_donor = load(f"{root}/qat/ViT_B_16__laion2b_s34b_b88k/EuroSAT/*/mult=1/qat=*/seed=2038/epoch_*.pt")
fp_recv = load(f"{root}/fp/ViT_B_16__laion2b_s34b_b88k/DTD/*/mult=1/seed=2038/epoch_*.pt")
lam = 1.0
patched = {k: fp_recv[k] + lam * (qat_donor[k] - fp_donor[k]) for k in fp_recv}
Load patched into the receiver model, apply 3-bit per-channel PTQ to every linear layer except the head, and evaluate. code/experiments/*/001_qat_transfer/qv_transfer.py in the code repository does exactly this.
Research graph
The questions behind each experiment phase, how they depend on one another and on the paper's propositions, and the scripts that answer them are laid out as a public Flywheel graph. Each node names its question, its method and the code in the code repository that implements it.
License
Finetuned from OpenCLIP LAION-2B checkpoints (MIT). The finetuning datasets carry their own terms.
Citation
@inproceedings{solombrino2026zeroshot,
title = {Zero-Shot Quantization via Weight-Space Arithmetic},
author = {Solombrino, Daniele and Gargiulo, Antonio Andrea and Zirilli, Alessandro and Zhou, Luca and Minut, Adrian Robert and Rodol{\`a}, Emanuele},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
url = {https://openreview.net/forum?id=wrUEnnSgXa}
}