FGIR-Backbones: backbones × training and evaluation settings on 17 fine-grained datasets

These are the checkpoints behind A Large-Scale Study on the Accuracy vs Cost Trade-offs of Training and Evaluation Settings in Fine-Grained Image Recognition (arXiv:2605.18700), presented at the FGVC13 workshop at CVPR 2026. The study trains 16 ImageNet-pretrained backbones (CNNs and vision transformers) on 17 fine-grained image recognition (FGIR) datasets under four training strategies, plus a 448 px set of runs, and compares accuracy against training and inference cost. Training code: arkel23/FGIR-Backbones. Loader: arkel23/fgir-zoo.

825 checkpoints, one per (dataset, backbone, strategy, image size), each from one training seed. The collection FGIR-Backbones (FGVC13 @ CVPR 2026) groups this repo with the paper.

Strategies

Strategy Meaning
ft full fine-tune
fz frozen backbone, linear head
cal CAL (counterfactual attention learning)
cal_cm CALMix (CAL with cross-image discriminative-region mixing)

CAL-NC and CALMix-NC, the paper's no-crop evaluation variants, are not separate files. They evaluate a cal or cal_cm checkpoint without the second forward pass on the attention-guided crop (cal_ap_only=True in the config).

Layout

One folder per strategy and image size, files named {dataset}_{model_name}_{strategy}.pth. manifest.csv lists every file with its dataset, backbone, strategy, image size, training serial and seed, class count, accuracy, SHA-256 and size.

Folder Strategy Image size Serial Backbones Datasets Files
ft_224 ft 224 1, 5 16 17 181
fz_224 fz 224 1, 5 16 17 181
cal_224 cal 224 1, 5 16 17 181
cal_cm_224 cal_cm 224 8 16 4 64
ft_384 ft 384 3, 6 16 4 64
fz_384 fz 384 3, 6 16 4 64
cal_384 cal 384 3, 6 16 4 64
cal_448 cal 448 15 5 3 13
cal_cm_448 cal_cm 448 11 5 3 13

Serial is the run group in the paper's experiment log. The 9 core backbones (VGG-19, ResNet-101, ResNetV2-101, BiT-M ResNetV2-101x3, ViT-B/16, BEiTv2-B/16, Swin-B, ConvNeXt-B, VAN-B3) are trained on all 17 datasets at 224 px; the other 7 on aircraft, cub, soygene and soylocal. At 384 px the Swin backbones use their 384 px variants (swin_*_window12_384*). The 448 px folders hold ViT-B/16 and four torchvision ResNets (resnet18, tv_resnet34, tv_resnet50, tv_resnet101) on aircraft, cars and cub.

Each file is a torch.save dict with four keys: config (the full training configuration, an argparse.Namespace), model (the state dict), accuracy and epoch. There is no optimizer state. The config drives the rebuild, so a file loads without the training repository.

Load a checkpoint and classify an image

fgir_zoo holds a frozen copy of the model code and pins timm==0.9.12.

import torch
from PIL import Image
from torchvision import transforms
from fgir_zoo import backbones

model = backbones.create_model('cal_224/cub_vit_b16_cal')  # or model_name='vit_b16', dataset='cub', strategy='cal'
cfg = model.config
tf = transforms.Compose([
    transforms.Resize((cfg.test_resize_size, cfg.test_resize_size),
                      interpolation=transforms.InterpolationMode.BICUBIC),
    transforms.CenterCrop(cfg.image_size),
    transforms.ToTensor(),
    transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
# a CUB-200-2011 test image, class index 50 (051.Horned_Grebe)
x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0)
with torch.no_grad():
    logits, _ = model(x)          # CAL returns (logits, attention crops)
    model.model.ap_only = True    # CAL-NC: skip the forward pass on the crop
    logits_nc = model(x)
print(logits.argmax(-1).item(), logits.softmax(-1).max().item())        # 50 0.9896
print(logits_nc.argmax(-1).item(), logits_nc.softmax(-1).max().item())  # 50 0.9902

The raw file is one Hub call away:

import torch
from huggingface_hub import hf_hub_download

path = hf_hub_download('ERISLab/FGIR-Backbones', 'cal_224/cub_vit_b16_cal.pth')
ckpt = torch.load(path, map_location='cpu', weights_only=False)
ckpt['config'].model_name, ckpt['accuracy']

Accuracy of the released checkpoints

Top-1 accuracy (%) stored in each file: the run's final accuracy on the dataset's test split, as recorded in the experiment log. Each value is one seed; the paper reports means over two or three seeds, so its tables differ slightly. The 17 files below 10% are all BEiTv2 or soyglobal (1,938 classes) cells that fail to train on every seed; they are the study's results, not damaged files.

ft_224: full fine-tune, 224 px

Backbone aircraft cars cotton cub dafb dogs flowers food inat17 moe nabirds pets soyageing soygene soyglobal soylocal vegfru
beitv2_base_patch16_224_in22k 85.33 88.84 2.92 86.95 92.56 88.87 98.59 91.49 71.54 94.05 87.90 77.65 43.31 9.42 2.94 14.50 95.95
convnext_base 86.35 81.96 30.32 19.50
convnext_base_in22k 87.49 80.10 37.08 84.52 87.92 86.77 99.46 90.65 67.40 93.24 83.84 92.94 35.86 36.88 2.89 26.17 94.41
convnext_large_in22k 83.74 88.30 33.67 33.83
deit3_base_patch16_224 84.79 81.86 45.42 24.50
deit3_base_patch16_224_in21ft1k 84.82 87.76 22.93 27.50
deit3_large_patch16_224_in21ft1k 85.99 90.30 49.39 20.00
resnet101 81.85 85.97 12.50 76.89 86.52 91.06 90.91 84.49 59.22 89.16 74.86 92.64 53.43 26.64 7.28 16.17 85.07
resnetv2_101 82.60 86.94 16.25 76.91 89.37 90.31 92.36 85.32 59.79 92.80 74.38 92.04 52.93 30.81 8.26 13.67 86.30
resnetv2_101x3_bitm_in21k 86.08 88.57 44.17 88.78 91.56 88.86 99.37 90.61 69.90 95.45 86.20 93.95 68.93 50.79 23.05 38.00 95.69
swin_base_patch4_window7_224 85.42 84.47 25.71 24.67
swin_base_patch4_window7_224_in22k 87.73 90.40 49.17 90.56 92.35 88.22 99.63 92.37 73.54 95.72 88.86 94.74 50.97 31.80 7.41 26.33 96.11
swin_large_patch4_window7_224_in22k 88.45 91.15 36.64 30.33
van_b3 86.47 89.39 41.67 79.32 91.26 95.05 96.05 88.46 64.61 94.67 82.50 95.07 65.54 37.49 8.34 20.83 90.27
vgg19_bn 78.55 86.56 30.42 76.13 85.75 85.58 94.28 81.71 53.77 91.68 72.06 92.29 43.78 53.92 21.26 31.67 82.38
vit_b16 82.12 86.69 50.42 87.83 89.96 90.89 99.32 89.66 66.06 94.90 83.77 93.89 37.25 31.16 18.78 25.83 93.65

fz_224: frozen backbone, linear head, 224 px

Backbone aircraft cars cotton cub dafb dogs flowers food inat17 moe nabirds pets soyageing soygene soyglobal soylocal vegfru
beitv2_base_patch16_224_in22k 55.36 64.36 29.58 90.49 41.45 89.95 99.48 90.38 66.58 73.80 87.18 93.73 26.16 15.92 6.66 22.50 94.86
convnext_base 53.92 66.36 26.42 17.50
convnext_base_in22k 62.20 67.43 40.00 87.73 55.28 89.07 99.53 90.00 61.27 84.54 81.24 93.89 38.99 32.52 14.93 36.33 95.21
convnext_large_in22k 62.89 89.28 36.77 37.17
deit3_base_patch16_224 61.36 69.81 23.68 21.83
deit3_base_patch16_224_in21ft1k 65.92 80.01 23.82 30.00
deit3_large_patch16_224_in21ft1k 70.87 82.34 26.61 26.17
resnet101 44.07 46.75 20.83 62.96 36.08 87.06 83.74 57.45 28.52 68.84 51.90 90.46 24.12 13.63 6.45 12.83 66.72
resnetv2_101 46.47 47.59 26.25 58.72 32.05 84.59 84.34 58.00 26.34 75.23 46.32 89.29 21.82 16.57 6.91 23.33 64.83
resnetv2_101x3_bitm_in21k 52.15 62.98 37.92 87.26 50.53 89.35 99.27 86.52 56.47 82.40 80.69 92.94 41.90 28.41 16.24 27.50 93.85
swin_base_patch4_window7_224 59.92 76.11 29.97 27.00
swin_base_patch4_window7_224_in22k 68.02 76.20 40.83 91.04 59.38 88.67 99.63 91.11 67.47 86.75 87.48 94.33 38.16 36.60 21.67 35.00 96.08
swin_large_patch4_window7_224_in22k 67.75 90.61 39.50 39.67
van_b3 56.02 61.85 36.25 70.02 41.19 95.34 92.29 70.04 31.61 80.12 58.61 92.70 32.24 21.92 10.47 24.00 76.92
vgg19_bn 47.58 47.28 28.33 62.00 31.83 85.45 86.66 60.66 29.33 76.08 51.45 89.48 32.18 20.68 10.32 24.33 71.06
vit_b16 58.21 62.37 31.67 86.19 65.44 87.75 98.70 82.78 53.24 88.21 79.13 91.47 40.55 17.31 13.52 23.17 90.24

cal_224: CAL (counterfactual attention learning), 224 px

Backbone aircraft cars cotton cub dafb dogs flowers food inat17 moe nabirds pets soyageing soygene soyglobal soylocal vegfru
beitv2_base_patch16_224_in22k 90.82 92.14 16.67 89.39 94.88 83.29 94.54 92.63 75.43 96.50 89.75 95.04 71.52 69.51 41.38 34.00 93.87
convnext_base 90.46 87.50 64.35 14.67
convnext_base_in22k 93.04 94.44 57.50 91.53 94.67 90.15 99.50 92.76 75.53 96.84 90.59 94.96 82.79 69.12 33.83 36.00 95.90
convnext_large_in22k 92.62 91.47 69.09 37.17
deit3_base_patch16_224 90.40 87.28 73.61 32.50
deit3_base_patch16_224_in21ft1k 91.06 89.18 54.68 34.17
deit3_large_patch16_224_in21ft1k 92.62 91.34 73.77 34.50
resnet101 82.21 89.99 35.42 85.85 89.97 92.17 96.83 87.11 62.52 95.07 84.94 93.92 54.02 31.86 23.89 27.17 91.41
resnetv2_101 81.82 91.74 30.83 85.66 89.41 91.18 96.42 86.85 62.44 95.14 85.25 93.32 48.91 31.23 21.50 25.17 90.47
resnetv2_101x3_bitm_in21k 91.69 94.02 42.92 89.73 94.66 88.68 99.24 90.54 72.25 95.07 88.34 93.89 57.54 72.99 32.97 36.50 94.28
swin_base_patch4_window7_224 91.48 86.40 69.69 20.83
swin_base_patch4_window7_224_in22k 91.96 94.27 42.92 90.82 94.12 88.72 99.63 92.73 76.02 97.04 90.14 95.31 54.67 76.70 50.65 38.00 95.98
swin_large_patch4_window7_224_in22k 92.59 91.51 76.14 39.83
van_b3 92.47 94.49 49.58 88.16 94.78 95.44 97.84 90.32 72.25 96.87 87.61 95.18 69.70 64.96 23.94 27.00 93.26
vgg19_bn 88.75 93.47 42.92 84.31 92.33 86.38 98.08 87.83 62.85 94.90 84.47 92.59 66.65 70.02 20.81 27.17 91.14
vit_b16 85.45 90.55 34.17 88.57 91.82 91.62 99.28 90.65 68.25 96.13 86.31 94.44 36.38 52.10 21.09 25.00 94.19

cal_cm_224: CALMix (CAL with cross-image discriminative-region mixing), 224 px

Backbone aircraft cub soygene soylocal
beitv2_base_patch16_224_in22k 91.90 91.08 71.72 39.33
convnext_base 91.72 89.26 62.55 41.17
convnext_base_in22k 92.95 92.03 67.65 45.33
convnext_large_in22k 93.40 91.97 70.21 48.67
deit3_base_patch16_224 90.73 87.87 75.62 55.50
deit3_base_patch16_224_in21ft1k 91.15 89.35 66.65 47.17
deit3_large_patch16_224_in21ft1k 93.19 91.42 77.47 52.00
resnet101 86.41 87.54 39.00 24.83
resnetv2_101 85.33 86.57 34.09 26.33
resnetv2_101x3_bitm_in21k 92.02 89.75 70.54 40.83
swin_base_patch4_window7_224 90.91 87.38 67.18 46.33
swin_base_patch4_window7_224_in22k 92.68 91.06 77.65 50.50
swin_large_patch4_window7_224_in22k 92.50 91.44 76.43 52.17
van_b3 92.74 88.94 63.48 37.17
vgg19_bn 92.50 87.25 68.90 45.67
vit_b16 85.87 89.77 58.66 39.67

ft_384: full fine-tune, 384 px

Backbone aircraft cub soygene soylocal
beitv2_base_patch16_224_in22k 76.66 81.57 7.43 9.33
convnext_base 89.50 83.24 40.16 22.17
convnext_base_in22k 88.57 78.46 40.47 28.67
convnext_large_in22k 84.13 87.66 51.49 32.67
deit3_base_patch16_224 87.04 84.52 62.12 34.50
deit3_base_patch16_224_in21ft1k 87.40 87.69 29.98 31.67
deit3_large_patch16_224_in21ft1k 88.27 91.13 60.82 31.00
resnet101 86.65 80.22 37.72 13.67
resnetv2_101 85.60 80.05 40.22 16.83
resnetv2_101x3_bitm_in21k 89.32 89.77 60.61 42.83
swin_base_patch4_window12_384 88.06 85.35 30.81 25.50
swin_base_patch4_window12_384_in22k 90.31 91.09 44.20 30.83
swin_large_patch4_window12_384_in22k 89.95 92.15 40.74 31.33
van_b3 88.36 75.09 38.92 22.83
vgg19_bn 79.78 75.49 58.06 34.33
vit_b16 85.84 88.99 46.00 25.50

fz_384: frozen backbone, linear head, 384 px

Backbone aircraft cub soygene soylocal
beitv2_base_patch16_224_in22k 9.03 5.82 1.40 7.67
convnext_base 50.53 57.94 34.19 20.67
convnext_base_in22k 61.63 82.31 39.35 41.33
convnext_large_in22k 61.27 85.83 41.45 40.83
deit3_base_patch16_224 64.00 77.94 36.43 25.50
deit3_base_patch16_224_in21ft1k 67.48 85.57 33.52 29.00
deit3_large_patch16_224_in21ft1k 70.99 87.25 31.80 31.00
resnet101 50.65 66.00 14.43 10.33
resnetv2_101 52.66 60.61 25.80 25.00
resnetv2_101x3_bitm_in21k 57.76 89.25 35.16 30.50
swin_base_patch4_window12_384 62.92 76.96 36.42 30.00
swin_base_patch4_window12_384_in22k 72.10 91.37 41.86 35.50
swin_large_patch4_window12_384_in22k 70.42 90.59 46.40 43.50
van_b3 53.92 58.70 24.25 23.67
vgg19_bn 55.36 62.63 30.77 23.67
vit_b16 49.83 85.52 15.61 22.00

cal_384: CAL (counterfactual attention learning), 384 px

Backbone aircraft cub soygene soylocal
beitv2_base_patch16_224_in22k 86.77 84.95 75.78 39.50
convnext_base 91.90 90.14 77.50 20.83
convnext_base_in22k 94.87 92.23 80.96 41.33
convnext_large_in22k 93.94 92.46 80.64 42.17
deit3_base_patch16_224 92.65 88.38 80.98 41.00
deit3_base_patch16_224_in21ft1k 93.13 89.56 67.06 39.83
deit3_large_patch16_224_in21ft1k 94.09 91.16 82.95 43.17
resnet101 87.76 89.09 51.21 31.00
resnetv2_101 87.46 88.16 51.13 24.17
resnetv2_101x3_bitm_in21k 92.35 90.87 78.52 35.83
swin_base_patch4_window12_384 93.13 87.61 78.36 19.67
swin_base_patch4_window12_384_in22k 93.49 91.85 83.16 40.00
swin_large_patch4_window12_384_in22k 94.09 91.73 82.78 42.50
van_b3 93.46 89.40 76.44 20.50
vgg19_bn 89.89 86.87 78.43 28.67
vit_b16 89.32 89.80 68.20 14.33

cal_448: CAL (counterfactual attention learning), 448 px

Backbone aircraft cars cub
resnet18 92.41 94.19 87.47
tv_resnet101 94.63 95.01 89.92
tv_resnet34 93.55 94.02 88.44
tv_resnet50 94.54 94.94 89.56
vit_b16 90.71

cal_cm_448: CALMix (CAL with cross-image discriminative-region mixing), 448 px

Backbone aircraft cars cub
resnet18 93.13 93.94 88.18
tv_resnet101 94.45 95.22 90.27
tv_resnet34 93.97 94.79 89.16
tv_resnet50 94.63 94.99 89.63
vit_b16 90.82

Requirements

  • fgir-zoo (pip install git+https://github.com/arkel23/fgir-zoo.git), which pins timm==0.9.12
  • torch (checked with 2.5.1)

Citation

@inproceedings{rios_large-scale_2026,
  title     = {A Large-Scale Study on the Accuracy vs Cost Trade-offs of Training and Evaluation
               Settings in Fine-Grained Image Recognition},
  author    = {Rios, Edwin Arkel and Surya, Augusto Christian and Gosal, Oswin and Mikael, Fernando and
               Nicole, Mary Madeline and Jang, Kisoon and Lai, Bo-Cheng and Hu, Min-Chun},
  booktitle = {The 13th Workshop on Fine-Grained Visual Categorization (FGVC13) at CVPR 2026},
  note      = {Non-archival extended abstract},
  year      = {2026},
  eprint    = {2605.18700},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi       = {10.48550/arXiv.2605.18700},
  url       = {https://arxiv.org/abs/2605.18700}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including ERISLab/FGIR-Backbones

Paper for ERISLab/FGIR-Backbones