ReCalCon / README.md
SoongE's picture
Initial upload
dcd36d6 verified
|
Raw History Blame Contribute Delete
2.9 kB
---
license: apache-2.0
base_model:
- openai/clip-vit-base-patch16
- openai/clip-vit-base-patch32
pipeline_tag: zero-shot-image-classification
tags:
- clip
- vision-language
- prompt-learning
- contrastive-learning
- distribution-shift
---
# ReCalCon
Checkpoints for **Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language
Models** (BMVC 2026). [Paper](https://arxiv.org/abs/2609.06967) · [Code](https://github.com/SoongE/ReCalCon)
Each folder holds `model.safetensors` (the full model state dict, fp32) and `config.json`, which names the released
config to evaluate it with.
## Usage
```bash
hf download SoongE/ReCalCon --include "distribution_shift/imagenet_b16/*" --local-dir checkpoints
python -m scripts.eval --config imagenet_b16 \
--eval.checkpoint checkpoints/distribution_shift/imagenet_b16/model.safetensors
```
## Checkpoints
**Distribution shift** (top-1 accuracy, %)
| Checkpoint | Row | ImageNet | -R | -A | -V2 | -Sketch | ObjectNet |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `distribution_shift/imagenet_b16` | ViT-B/16, Update, ImageNet | 83.284 | 71.607 | 54.893 | 74.1 | 51.781 | 58.388 |
| `distribution_shift/imagenet_b32` | ViT-B/32, Update, ImageNet | 79.354 | 62.41 | 32.573 | 68.15 | 44.503 | 49.203 |
| `distribution_shift/imagenet_petl_b16` | ViT-B/16, Freeze (PETL), ImageNet | 80.206 | 73.163 | 50.867 | 70.13 | 49.614 | 55.691 |
| `distribution_shift/imagenet_petl_b32` | ViT-B/32, Freeze (PETL), ImageNet | 76.12 | 63.563 | 30.453 | 64.52 | 42.267 | 46.662 |
iWildCam (macro-F1, %). The released evaluation reports the OOD split.
| Checkpoint | Row | ID (macro-F1) | OOD (macro-F1) |
| --- | --- | --- | --- |
| `distribution_shift/iwildcam_b16` | ViT-B/16, Update, iWildCam | 51.253 | 37.803 |
| `distribution_shift/iwildcam_b32` | ViT-B/32, Update, iWildCam | 42.116 | 29.275 |
| `distribution_shift/iwildcam_petl_b16` | ViT-B/16, Freeze (PETL), iWildCam | 45.49 | 30.327 |
| `distribution_shift/iwildcam_petl_b32` | ViT-B/32, Freeze (PETL), iWildCam | 34.709 | 24.502 |
**Transfer learning** (top-1 accuracy, %)
| Checkpoint | Dataset | Accuracy |
| --- | --- | --- |
| `transfer/caltech101` | Caltech101 | 97.85 |
| `transfer/flowers102` | Flowers102 | 98.894 |
| `transfer/pcam` | PCam | 89.413 |
| `transfer/stanfordcars` | StanfordCars | 91.419 |
## License
Apache License 2.0. The models are fine-tuned from OpenAI CLIP (MIT License); each evaluation dataset keeps its own
terms.
## Citation
```bibtex
@inproceedings{oh2026recalibrated,
title = {Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models},
author = {Oh, Seungmin and Kang, Seunghun and Ryu, Jongbin},
booktitle = {Proceedings of the British Machine Vision Conference},
year = {2026},
publisher = {BMVA},
url = {https://arxiv.org/abs/2609.06967}
}
```