File size: 2,901 Bytes
dcd36d6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 | ---
license: apache-2.0
base_model:
- openai/clip-vit-base-patch16
- openai/clip-vit-base-patch32
pipeline_tag: zero-shot-image-classification
tags:
- clip
- vision-language
- prompt-learning
- contrastive-learning
- distribution-shift
---
# ReCalCon
Checkpoints for **Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language
Models** (BMVC 2026). [Paper](https://arxiv.org/abs/2609.06967) · [Code](https://github.com/SoongE/ReCalCon)
Each folder holds `model.safetensors` (the full model state dict, fp32) and `config.json`, which names the released
config to evaluate it with.
## Usage
```bash
hf download SoongE/ReCalCon --include "distribution_shift/imagenet_b16/*" --local-dir checkpoints
python -m scripts.eval --config imagenet_b16 \
--eval.checkpoint checkpoints/distribution_shift/imagenet_b16/model.safetensors
```
## Checkpoints
**Distribution shift** (top-1 accuracy, %)
| Checkpoint | Row | ImageNet | -R | -A | -V2 | -Sketch | ObjectNet |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `distribution_shift/imagenet_b16` | ViT-B/16, Update, ImageNet | 83.284 | 71.607 | 54.893 | 74.1 | 51.781 | 58.388 |
| `distribution_shift/imagenet_b32` | ViT-B/32, Update, ImageNet | 79.354 | 62.41 | 32.573 | 68.15 | 44.503 | 49.203 |
| `distribution_shift/imagenet_petl_b16` | ViT-B/16, Freeze (PETL), ImageNet | 80.206 | 73.163 | 50.867 | 70.13 | 49.614 | 55.691 |
| `distribution_shift/imagenet_petl_b32` | ViT-B/32, Freeze (PETL), ImageNet | 76.12 | 63.563 | 30.453 | 64.52 | 42.267 | 46.662 |
iWildCam (macro-F1, %). The released evaluation reports the OOD split.
| Checkpoint | Row | ID (macro-F1) | OOD (macro-F1) |
| --- | --- | --- | --- |
| `distribution_shift/iwildcam_b16` | ViT-B/16, Update, iWildCam | 51.253 | 37.803 |
| `distribution_shift/iwildcam_b32` | ViT-B/32, Update, iWildCam | 42.116 | 29.275 |
| `distribution_shift/iwildcam_petl_b16` | ViT-B/16, Freeze (PETL), iWildCam | 45.49 | 30.327 |
| `distribution_shift/iwildcam_petl_b32` | ViT-B/32, Freeze (PETL), iWildCam | 34.709 | 24.502 |
**Transfer learning** (top-1 accuracy, %)
| Checkpoint | Dataset | Accuracy |
| --- | --- | --- |
| `transfer/caltech101` | Caltech101 | 97.85 |
| `transfer/flowers102` | Flowers102 | 98.894 |
| `transfer/pcam` | PCam | 89.413 |
| `transfer/stanfordcars` | StanfordCars | 91.419 |
## License
Apache License 2.0. The models are fine-tuned from OpenAI CLIP (MIT License); each evaluation dataset keeps its own
terms.
## Citation
```bibtex
@inproceedings{oh2026recalibrated,
title = {Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models},
author = {Oh, Seungmin and Kang, Seunghun and Ryu, Jongbin},
booktitle = {Proceedings of the British Machine Vision Conference},
year = {2026},
publisher = {BMVA},
url = {https://arxiv.org/abs/2609.06967}
}
```
|