Instructions to use ERISLab/Hierarchical-Backbones with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use ERISLab/Hierarchical-Backbones with timm:
import timm model = timm.create_model("hf_hub:ERISLab/Hierarchical-Backbones", pretrained=True) - Notebooks
- Google Colab
- Kaggle
Hierarchical-Backbones: hierarchical classification across pretrained backbones
These are the checkpoints behind Revisiting the Backbone, Pretraining and Transferability for Hierarchical Fine-Grained Image Recognition, CVGIP 2025 (arXiv link to be added). Twenty-one pretrained backbones (ResNet-50, ViT-B, DeiT-B and DeiT-3-B under supervised, self-supervised and vision-language pretraining) fine-tuned for fine-grained recognition with a hierarchical classification head. The head predicts every level of the annotated label hierarchy, coarsest first: three levels on CUB and Aircraft, two on Cars. Code: arkel23/Hierarchical.
168 checkpoints, one per configuration, each the last epoch of one training run at 448 px.
Each file is a torch.save dict with config (the full training configuration), model (the
state dict of the backbone and head), accuracy and epoch, with no optimizer state. File names
are the runs' experiment-log names: dataset, cluster ratio for a pseudo-hierarchy, model, fz for
a frozen backbone, and the serial. Load them with
fgir-zoo. The collection is this work's
page on the Hub; the paper joins it once it is on arXiv.
Layout
One folder per serial.
| Folder | What | Files | Mean accuracy |
|---|---|---|---|
serial_21 |
Real (annotated) hierarchy, full fine-tuning, CUB/Aircraft/Cars x 21 backbones, 200 epochs, each backbone at its best learning rate from the paper's sweep, seed 1 | 63 | 90.39 |
serial_58 |
Aircraft x 21 backbones, 50 epochs, best learning rate, with the real hierarchy; matched with serial_59 |
21 | 91.39 |
serial_59 |
Aircraft x 21 backbones, 50 epochs, best learning rate, without a hierarchy; matched with serial_58 |
21 | 90.94 |
serial_60 |
Real hierarchy with a frozen backbone (only the classification head trained), the same 63 configurations at their best learning rate, 50 or 200 epochs | 63 | 64.90 |
manifest.csv lists every file with its dataset, model, hierarchy (real, pseudo or none),
cluster ratio, frozen backbone, serial, seed, epochs, image size, class count, number of levels,
accuracy, SHA-256 and size.
Load a checkpoint and classify an image
With a hierarchy the model returns a tuple of logits, one per level from coarsest to finest; without one it returns a single logits tensor. The finest level is the prediction. The evaluation transform is the training code's: a bicubic resize to 550 x 550, a 448 px center crop and ImageNet normalization.
import torch
from PIL import Image
from torchvision import transforms
from fgir_zoo import hierarchical
model = hierarchical.create_model('serial_21/cub_hideit_base_patch16_224.fb_in1k_21')
cfg = model.config
tf = transforms.Compose([
transforms.Resize((cfg.test_resize_size, cfg.test_resize_size),
interpolation=transforms.InterpolationMode.BICUBIC),
transforms.CenterCrop(cfg.image_size),
transforms.ToTensor(),
transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
# a CUB-200-2011 test image, class index 50 (051.Horned_Grebe)
x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0)
with torch.no_grad():
order, family, species = model(x) # one logits tensor per level, coarsest first
print(species.argmax(-1).item(), species.softmax(-1).max().item()) # 50 0.8240
Accuracy of the released checkpoints
Top-1 accuracy (%) stored in each file: the run's own finest-level accuracy on the dataset's test
split after the last epoch. The papers report the maximum over a learning-rate sweep, so their
tables differ. Per-file accuracy is in manifest.csv; the table below covers
serial_21.
| Model | aircraft (21) | cars (21) | cub (21) |
|---|---|---|---|
hideit3_base_patch16_224.fb_in1k |
92.08 | 93.40 | 88.63 |
hideit3_base_patch16_224.fb_in22k_ft_in1k |
92.98 | 93.65 | 90.63 |
hideit_base_patch16_224.fb_in1k |
92.11 | 93.06 | 88.82 |
hiresnet50.a1_in1k |
92.92 | 93.14 | 85.78 |
hiresnet50.fb_ssl_yfcc100m_ft_in1k |
93.07 | 94.28 | 85.55 |
hiresnet50.fb_swsl_ig1b_ft_in1k |
93.10 | 94.28 | 85.48 |
hiresnet50.gluon_in1k |
93.01 | 94.23 | 85.24 |
hiresnet50.in1k_mocov3 |
93.07 | 93.53 | 84.85 |
hiresnet50.in1k_spark |
91.27 | 92.38 | 78.63 |
hiresnet50.in1k_supcon |
93.34 | 94.24 | 85.38 |
hiresnet50.in1k_swav |
92.65 | 93.86 | 86.02 |
hiresnet50.in21k_miil |
93.01 | 93.74 | 87.85 |
hiresnet50.tv2_in1k |
92.80 | 93.81 | 87.28 |
hiresnet50.tv_in1k |
92.71 | 94.09 | 86.16 |
hivit_base_patch16_224.dino |
88.00 | 91.10 | 86.47 |
hivit_base_patch16_224.in1k_mocov3 |
90.19 | 93.07 | 87.73 |
hivit_base_patch16_224.mae |
83.62 | 92.79 | 85.26 |
hivit_base_patch16_224.orig_in21k |
89.32 | 91.73 | 90.97 |
hivit_base_patch16_224_miil.in21k |
90.49 | 92.13 | 90.21 |
hivit_base_patch16_clip_224.laion2b |
86.41 | 92.36 | 80.45 |
hivit_base_patch16_siglip_224.v2_webli |
93.40 | 95.06 | 87.81 |
Requirements
fgir-zoo(pip install git+https://github.com/arkel23/fgir-zoo.git), which pinstimm==0.9.12torch(checked with 2.5.1)
Citation
@inproceedings{surya_revisiting_2025,
title = {Revisiting the Backbone, Pretraining and Transferability for Hierarchical
Fine-Grained Image Recognition},
author = {Surya, Augusto Christian and Rios, Edwin Arkel and Lai, Bo-Cheng and Hu, Min-Chun},
booktitle = {Computer Vision, Graphics, and Image Processing (CVGIP)},
year = {2025},
eprint = {TBA},
note = {A. C. Surya and E. A. Rios contributed equally. arXiv ID to be added.}
}