File size: 3,146 Bytes
0364d34
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a9468ec
0364d34
 
 
 
 
 
 
 
 
 
 
 
 
b54bf45
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
---
pipeline_tag: image-classification
library_name: fgir-zoo
tags:
- fine-grained
- image-classification
- vision-transformer
- full-fine-tuning
---

# FGIRFT

A full fine-tuning benchmark for fine-grained image recognition (FGIR). Every checkpoint is an
ImageNet-21k **ViT-B/16 fine-tuned end to end** — the whole backbone is trained, not frozen — with a
fine-grained classification head, across 21 fine-grained datasets. FGIRFT is the full fine-tuning
counterpart to the ERIS Lab's parameter-efficient works (the ILA/SAW and AAA adapters in
`ERISLab/FGIR-ViT`, which keep the backbone frozen and train only small adapters).

242 checkpoints, one seed per configuration, last epoch only, grouped by input resolution: `ft_224`
(111) and `ft_448` (131). Everything is described in `manifest.csv` and loadable by name with the
standalone `fgir_zoo` library (no training repo needed).

## What it holds

ViT-B/16 fully fine-tuned at 224 or 448 px, each with one FGIR classification head: a plain linear
classifier (`cls`), GLSim, MAWS (the FFVT head), PSM (the TransFG head), CAL, MPN-COV, and
attention-rollout heads. The 21 datasets span general fine-grained sets (CUB, NABirds, Aircraft,
Stanford Cars, Stanford Dogs, Oxford Flowers, Oxford Pets, Food-101, VegFru, DAFB, Moe) and the
ultra-fine-grained leaf sets (Cotton, SoyAgeing and its five single-season subsets, SoyGene,
SoyGlobal, SoyLocal). `num_classes` ranges from 37 to 3263. `manifest.csv` maps every file to its
dataset, head, serial, resolution, `num_classes`, top-1 accuracy and sha256.

## Load a checkpoint

Standalone, no `fgir_vit` install — the model code is vendored inside `fgir_zoo`:

```bash
pip install git+https://github.com/arkel23/fgir-zoo.git
```

```python
import fgir_zoo
model = fgir_zoo.load('fgirvit/ft_448/aircraft_vit_b16_avg_cls_rollout_3')
model.eval()
# fgir_zoo.list_models(family='fgirvit', group='ft_448') lists every name; model.config holds the run's settings.
```

Or fetch the file directly and load it yourself:

```python
import torch
from huggingface_hub import hf_hub_download
path = hf_hub_download('ERISLab/FGIRFT', 'ft_448/aircraft_vit_b16_avg_cls_rollout_1003.pth')
ckpt = torch.load(path, map_location='cpu', weights_only=False)   # dict: config, model, accuracy, epoch
```

## Requirements

- `torch>=2.5`, `timm==0.9.12` (the checkpoints were trained on it and the vendored model code forks
  that version), `huggingface_hub`, `safetensors`, `einops`, `ml_collections`.

## Related

Part of the [ERISLab FGIRFT collection](https://huggingface.co/collections/ERISLab/fgirft-full-fine-tuning-benchmark-for-fgir-6ab2e8940f3a28ec7ba0bfbf).
For the parameter-efficient adapter works (ILA, SAW, AAA) on a frozen ViT-B/16, see
[`ERISLab/FGIR-ViT`](https://huggingface.co/ERISLab/FGIR-ViT). As a full fine-tuning line, FGIRFT sits
close to the [ERISLab FGIR-Backbones benchmark](https://huggingface.co/collections/ERISLab/fgir-backbones-fgvc13-cvpr-2026-6ab234cc19fcfd2de1f796a4)
(repo [`ERISLab/FGIR-Backbones`](https://huggingface.co/ERISLab/FGIR-Backbones)), which studies
backbones and pretraining recipes for fine-grained recognition.