Hierarchical-Backbones: hierarchical classification across pretrained backbones

These are the checkpoints behind Revisiting the Backbone, Pretraining and Transferability for Hierarchical Fine-Grained Image Recognition, CVGIP 2025 (arXiv link to be added). Twenty-one pretrained backbones (ResNet-50, ViT-B, DeiT-B and DeiT-3-B under supervised, self-supervised and vision-language pretraining) fine-tuned for fine-grained recognition with a hierarchical classification head. The head predicts every level of the annotated label hierarchy, coarsest first: three levels on CUB and Aircraft, two on Cars. Code: arkel23/Hierarchical.

168 checkpoints, one per configuration, each the last epoch of one training run at 448 px. Each file is a torch.save dict with config (the full training configuration), model (the state dict of the backbone and head), accuracy and epoch, with no optimizer state. File names are the runs' experiment-log names: dataset, cluster ratio for a pseudo-hierarchy, model, fz for a frozen backbone, and the serial. Load them with fgir-zoo. The collection is this work's page on the Hub; the paper joins it once it is on arXiv.

Layout

One folder per serial.

Folder What Files Mean accuracy
serial_21 Real (annotated) hierarchy, full fine-tuning, CUB/Aircraft/Cars x 21 backbones, 200 epochs, each backbone at its best learning rate from the paper's sweep, seed 1 63 90.39
serial_58 Aircraft x 21 backbones, 50 epochs, best learning rate, with the real hierarchy; matched with serial_59 21 91.39
serial_59 Aircraft x 21 backbones, 50 epochs, best learning rate, without a hierarchy; matched with serial_58 21 90.94
serial_60 Real hierarchy with a frozen backbone (only the classification head trained), the same 63 configurations at their best learning rate, 50 or 200 epochs 63 64.90

manifest.csv lists every file with its dataset, model, hierarchy (real, pseudo or none), cluster ratio, frozen backbone, serial, seed, epochs, image size, class count, number of levels, accuracy, SHA-256 and size.

Load a checkpoint and classify an image

With a hierarchy the model returns a tuple of logits, one per level from coarsest to finest; without one it returns a single logits tensor. The finest level is the prediction. The evaluation transform is the training code's: a bicubic resize to 550 x 550, a 448 px center crop and ImageNet normalization.

import torch
from PIL import Image
from torchvision import transforms
from fgir_zoo import hierarchical

model = hierarchical.create_model('serial_21/cub_hideit_base_patch16_224.fb_in1k_21')
cfg = model.config
tf = transforms.Compose([
    transforms.Resize((cfg.test_resize_size, cfg.test_resize_size),
                      interpolation=transforms.InterpolationMode.BICUBIC),
    transforms.CenterCrop(cfg.image_size),
    transforms.ToTensor(),
    transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
# a CUB-200-2011 test image, class index 50 (051.Horned_Grebe)
x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0)
with torch.no_grad():
    order, family, species = model(x)  # one logits tensor per level, coarsest first
print(species.argmax(-1).item(), species.softmax(-1).max().item())  # 50 0.8240

Accuracy of the released checkpoints

Top-1 accuracy (%) stored in each file: the run's own finest-level accuracy on the dataset's test split after the last epoch. The papers report the maximum over a learning-rate sweep, so their tables differ. Per-file accuracy is in manifest.csv; the table below covers serial_21.

Model aircraft (21) cars (21) cub (21)
hideit3_base_patch16_224.fb_in1k 92.08 93.40 88.63
hideit3_base_patch16_224.fb_in22k_ft_in1k 92.98 93.65 90.63
hideit_base_patch16_224.fb_in1k 92.11 93.06 88.82
hiresnet50.a1_in1k 92.92 93.14 85.78
hiresnet50.fb_ssl_yfcc100m_ft_in1k 93.07 94.28 85.55
hiresnet50.fb_swsl_ig1b_ft_in1k 93.10 94.28 85.48
hiresnet50.gluon_in1k 93.01 94.23 85.24
hiresnet50.in1k_mocov3 93.07 93.53 84.85
hiresnet50.in1k_spark 91.27 92.38 78.63
hiresnet50.in1k_supcon 93.34 94.24 85.38
hiresnet50.in1k_swav 92.65 93.86 86.02
hiresnet50.in21k_miil 93.01 93.74 87.85
hiresnet50.tv2_in1k 92.80 93.81 87.28
hiresnet50.tv_in1k 92.71 94.09 86.16
hivit_base_patch16_224.dino 88.00 91.10 86.47
hivit_base_patch16_224.in1k_mocov3 90.19 93.07 87.73
hivit_base_patch16_224.mae 83.62 92.79 85.26
hivit_base_patch16_224.orig_in21k 89.32 91.73 90.97
hivit_base_patch16_224_miil.in21k 90.49 92.13 90.21
hivit_base_patch16_clip_224.laion2b 86.41 92.36 80.45
hivit_base_patch16_siglip_224.v2_webli 93.40 95.06 87.81

Requirements

  • fgir-zoo (pip install git+https://github.com/arkel23/fgir-zoo.git), which pins timm==0.9.12
  • torch (checked with 2.5.1)

Citation

@inproceedings{surya_revisiting_2025,
  title     = {Revisiting the Backbone, Pretraining and Transferability for Hierarchical
               Fine-Grained Image Recognition},
  author    = {Surya, Augusto Christian and Rios, Edwin Arkel and Lai, Bo-Cheng and Hu, Min-Chun},
  booktitle = {Computer Vision, Graphics, and Image Processing (CVGIP)},
  year      = {2025},
  eprint    = {TBA},
  note      = {A. C. Surya and E. A. Rios contributed equally. arXiv ID to be added.}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including ERISLab/Hierarchical-Backbones