Instructions to use gladia/qat-transfer-text with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use gladia/qat-transfer-text with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="gladia/qat-transfer-text")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("gladia/qat-transfer-text", device_map="auto") - Notebooks
- Google Colab
- Kaggle
qat-transfer checkpoints: Text classifiers: BERT-base/large, EmbeddingGemma-300M, Qwen3-Embedding-0.6B
Finetuned checkpoints for Zero-Shot Quantization via Weight-Space Arithmetic (NeurIPS 2026). Code: github.com/gladia-research-group/qat-transfer · arXiv · OpenReview · Research graph
For every backbone and dataset this repo has a pair of checkpoints that share a key set exactly:
- FP: full-precision finetune.
- QAT: 3-bit, per-channel quantization-aware finetune (STE), saved with the QAT wrappers stripped.
Their difference is the quantization vector QV = QAT - FP. Adding a donor task's QV to a different receiver task's FP checkpoint, FP_receiver + lambda * QV_donor, recovers much of the receiver's QAT benefit under 3-bit post-training quantization, with no receiver-side QAT and, at lambda = 1, no receiver data.
The other model families are in gladia/qat-transfer-timm, gladia/qat-transfer-open_clip and gladia/qat-transfer-text.
Contents
Backbones: google-bert/bert-base-uncased, google-bert/bert-large-uncased, google/embeddinggemma-300m, Qwen/Qwen3-Embedding-0.6B.
Datasets: Emotion, IMDB, Banking77, AmazonReviewsClassification, AmazonCounterfactual, MassiveIntent, MassiveScenario, MTOPDomain, MTOPIntent, ToxicConversations, TweetSentimentExtraction.
Shared configuration: AdamW, lr=1e-5, wd=0.1, no label smoothing, gradient clipping at 1.0, seed 2038, the 1x training schedule (mult=1). QAT and evaluation-time PTQ both use 3 bits, per-channel, with the classification head (classifier (BERT) / score (Gemma, Qwen) left unquantized. Larger models were trained at a smaller per-device batch with gradient accumulation; the optim=` directory names the per-device batch size.
Each run directory holds backbone_epoch_{N}.pt (state dict without the head) and head_epoch_{N}.pt (head state dict). N is the dataset's reference epoch count, and run_meta.json, where present, records the realized schedule.
Layout
The repo root mirrors the code's storage/ directory, so the scripts in the code repository read these files with no changes:
checkpoints/text/ilharco_automodelforsequenceclassification/fp/{model}/{dataset}/optim=.../mult=1/seed=2038/
checkpoints/text/ilharco_automodelforsequenceclassification/qat/{model}/{dataset}/optim=.../mult=1/qat=bits=3_gran=channel_skip=.../seed=2038/
Download
pip install -U huggingface_hub
hf download gladia/qat-transfer-text --local-dir storage
Download one backbone only:
hf download gladia/qat-transfer-text --local-dir storage --include "*/google_bert_bert_base_uncased/*"
Then set CHECKPOINT_BASE_PATH=storage/checkpoints and HEAD_BASE_PATH=storage/heads in the code repository's .env. All three repos can be downloaded into the same storage/.
Building a quantization vector by hand
The checkpoints are plain PyTorch state dicts, so the method needs nothing beyond torch:
import glob, torch
def load(pattern):
(path,) = glob.glob(pattern)
return torch.load(path, map_location="cpu")
root = "storage/checkpoints/text/ilharco_automodelforsequenceclassification"
fp_donor = load(f"{root}/fp/google_bert_bert_base_uncased/Emotion/*/mult=1/seed=2038/backbone_epoch_*.pt")
qat_donor = load(f"{root}/qat/google_bert_bert_base_uncased/Emotion/*/mult=1/qat=*/seed=2038/backbone_epoch_*.pt")
fp_recv = load(f"{root}/fp/google_bert_bert_base_uncased/IMDB/*/mult=1/seed=2038/backbone_epoch_*.pt")
lam = 1.0
patched = {k: fp_recv[k] + lam * (qat_donor[k] - fp_donor[k]) for k in fp_recv}
Load patched into the receiver model, apply 3-bit per-channel PTQ to every linear layer except the head, and evaluate. code/experiments/*/001_qat_transfer/qv_transfer.py in the code repository does exactly this.
Research graph
The questions behind each experiment phase, how they depend on one another and on the paper's propositions, and the scripts that answer them are laid out as a public Flywheel graph. Each node names its question, its method and the code in the code repository that implements it.
License
Mixed upstream licenses. BERT (google-bert/*) and Qwen/Qwen3-Embedding-0.6B are Apache-2.0. google/embeddinggemma-300m finetunes are Gemma derivatives and are distributed under the Gemma Terms of Use, including its Prohibited Use Policy. By using those files you agree to those terms. The Gemma finetunes are modified versions of the original model; see NOTICE. The finetuning datasets carry their own terms.
Citation
@inproceedings{solombrino2026zeroshot,
title = {Zero-Shot Quantization via Weight-Space Arithmetic},
author = {Solombrino, Daniele and Gargiulo, Antonio Andrea and Zirilli, Alessandro and Zhou, Luca and Minut, Adrian Robert and Rodol{\`a}, Emanuele},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
url = {https://openreview.net/forum?id=wrUEnnSgXa}
}