File size: 2,901 Bytes
dcd36d6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
---
license: apache-2.0
base_model:
- openai/clip-vit-base-patch16
- openai/clip-vit-base-patch32
pipeline_tag: zero-shot-image-classification
tags:
- clip
- vision-language
- prompt-learning
- contrastive-learning
- distribution-shift
---

# ReCalCon

Checkpoints for **Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language
Models** (BMVC 2026). [Paper](https://arxiv.org/abs/2609.06967) · [Code](https://github.com/SoongE/ReCalCon)

Each folder holds `model.safetensors` (the full model state dict, fp32) and `config.json`, which names the released
config to evaluate it with.

## Usage

```bash
hf download SoongE/ReCalCon --include "distribution_shift/imagenet_b16/*" --local-dir checkpoints
python -m scripts.eval --config imagenet_b16 \
  --eval.checkpoint checkpoints/distribution_shift/imagenet_b16/model.safetensors
```

## Checkpoints

**Distribution shift** (top-1 accuracy, %)

| Checkpoint | Row | ImageNet | -R | -A | -V2 | -Sketch | ObjectNet |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `distribution_shift/imagenet_b16` | ViT-B/16, Update, ImageNet | 83.284 | 71.607 | 54.893 | 74.1 | 51.781 | 58.388 |
| `distribution_shift/imagenet_b32` | ViT-B/32, Update, ImageNet | 79.354 | 62.41 | 32.573 | 68.15 | 44.503 | 49.203 |
| `distribution_shift/imagenet_petl_b16` | ViT-B/16, Freeze (PETL), ImageNet | 80.206 | 73.163 | 50.867 | 70.13 | 49.614 | 55.691 |
| `distribution_shift/imagenet_petl_b32` | ViT-B/32, Freeze (PETL), ImageNet | 76.12 | 63.563 | 30.453 | 64.52 | 42.267 | 46.662 |

iWildCam (macro-F1, %). The released evaluation reports the OOD split.

| Checkpoint | Row | ID (macro-F1) | OOD (macro-F1) |
| --- | --- | --- | --- |
| `distribution_shift/iwildcam_b16` | ViT-B/16, Update, iWildCam | 51.253 | 37.803 |
| `distribution_shift/iwildcam_b32` | ViT-B/32, Update, iWildCam | 42.116 | 29.275 |
| `distribution_shift/iwildcam_petl_b16` | ViT-B/16, Freeze (PETL), iWildCam | 45.49 | 30.327 |
| `distribution_shift/iwildcam_petl_b32` | ViT-B/32, Freeze (PETL), iWildCam | 34.709 | 24.502 |

**Transfer learning** (top-1 accuracy, %)

| Checkpoint | Dataset | Accuracy |
| --- | --- | --- |
| `transfer/caltech101` | Caltech101 | 97.85 |
| `transfer/flowers102` | Flowers102 | 98.894 |
| `transfer/pcam` | PCam | 89.413 |
| `transfer/stanfordcars` | StanfordCars | 91.419 |

## License

Apache License 2.0. The models are fine-tuned from OpenAI CLIP (MIT License); each evaluation dataset keeps its own
terms.

## Citation

```bibtex
@inproceedings{oh2026recalibrated,
  title     = {Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models},
  author    = {Oh, Seungmin and Kang, Seunghun and Ryu, Jongbin},
  booktitle = {Proceedings of the British Machine Vision Conference},
  year      = {2026},
  publisher = {BMVA},
  url       = {https://arxiv.org/abs/2609.06967}
}
```