--- pipeline_tag: image-classification library_name: pytorch tags: - fine-grained-image-recognition - image-classification - hierarchical-classification - timm --- # Hierarchical-Backbones: hierarchical classification across pretrained backbones These are the checkpoints behind *Revisiting the Backbone, Pretraining and Transferability for Hierarchical Fine-Grained Image Recognition*, CVGIP 2025 (arXiv link to be added). Twenty-one pretrained backbones (ResNet-50, ViT-B, DeiT-B and DeiT-3-B under supervised, self-supervised and vision-language pretraining) fine-tuned for fine-grained recognition with a hierarchical classification head. The head predicts every level of the annotated label hierarchy, coarsest first: three levels on CUB and Aircraft, two on Cars. Code: [arkel23/Hierarchical](https://github.com/arkel23/Hierarchical). 168 checkpoints, one per configuration, each the last epoch of one training run at 448 px. Each file is a `torch.save` dict with `config` (the full training configuration), `model` (the state dict of the backbone and head), `accuracy` and `epoch`, with no optimizer state. File names are the runs' experiment-log names: dataset, cluster ratio for a pseudo-hierarchy, model, `fz` for a frozen backbone, and the serial. Load them with [fgir-zoo](https://github.com/arkel23/fgir-zoo). The [collection](https://huggingface.co/collections/ERISLab/hierarchical-fgir-backbones-and-pretraining-cvgip-2025-6abab163857f4e028b3a38a1) is this work's page on the Hub; the paper joins it once it is on arXiv. ## Layout One folder per serial. | Folder | What | Files | Mean accuracy | |---|---|---|---| | `serial_21` | Real (annotated) hierarchy, full fine-tuning, CUB/Aircraft/Cars x 21 backbones, 200 epochs, each backbone at its best learning rate from the paper's sweep, seed 1 | 63 | 90.39 | | `serial_58` | Aircraft x 21 backbones, 50 epochs, best learning rate, with the real hierarchy; matched with `serial_59` | 21 | 91.39 | | `serial_59` | Aircraft x 21 backbones, 50 epochs, best learning rate, without a hierarchy; matched with `serial_58` | 21 | 90.94 | | `serial_60` | Real hierarchy with a frozen backbone (only the classification head trained), the same 63 configurations at their best learning rate, 50 or 200 epochs | 63 | 64.90 | `manifest.csv` lists every file with its dataset, model, hierarchy (`real`, `pseudo` or `none`), cluster ratio, frozen backbone, serial, seed, epochs, image size, class count, number of levels, accuracy, SHA-256 and size. ## Load a checkpoint and classify an image With a hierarchy the model returns a tuple of logits, one per level from coarsest to finest; without one it returns a single logits tensor. The finest level is the prediction. The evaluation transform is the training code's: a bicubic resize to 550 x 550, a 448 px center crop and ImageNet normalization. ```python import torch from PIL import Image from torchvision import transforms from fgir_zoo import hierarchical model = hierarchical.create_model('serial_21/cub_hideit_base_patch16_224.fb_in1k_21') cfg = model.config tf = transforms.Compose([ transforms.Resize((cfg.test_resize_size, cfg.test_resize_size), interpolation=transforms.InterpolationMode.BICUBIC), transforms.CenterCrop(cfg.image_size), transforms.ToTensor(), transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)), ]) # a CUB-200-2011 test image, class index 50 (051.Horned_Grebe) x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0) with torch.no_grad(): order, family, species = model(x) # one logits tensor per level, coarsest first print(species.argmax(-1).item(), species.softmax(-1).max().item()) # 50 0.8240 ``` ## Accuracy of the released checkpoints Top-1 accuracy (%) stored in each file: the run's own finest-level accuracy on the dataset's test split after the last epoch. The papers report the maximum over a learning-rate sweep, so their tables differ. Per-file accuracy is in `manifest.csv`; the table below covers `serial_21`. | Model | aircraft (21) | cars (21) | cub (21) | |---|---|---|---| | `hideit3_base_patch16_224.fb_in1k` | 92.08 | 93.40 | 88.63 | | `hideit3_base_patch16_224.fb_in22k_ft_in1k` | 92.98 | 93.65 | 90.63 | | `hideit_base_patch16_224.fb_in1k` | 92.11 | 93.06 | 88.82 | | `hiresnet50.a1_in1k` | 92.92 | 93.14 | 85.78 | | `hiresnet50.fb_ssl_yfcc100m_ft_in1k` | 93.07 | 94.28 | 85.55 | | `hiresnet50.fb_swsl_ig1b_ft_in1k` | 93.10 | 94.28 | 85.48 | | `hiresnet50.gluon_in1k` | 93.01 | 94.23 | 85.24 | | `hiresnet50.in1k_mocov3` | 93.07 | 93.53 | 84.85 | | `hiresnet50.in1k_spark` | 91.27 | 92.38 | 78.63 | | `hiresnet50.in1k_supcon` | 93.34 | 94.24 | 85.38 | | `hiresnet50.in1k_swav` | 92.65 | 93.86 | 86.02 | | `hiresnet50.in21k_miil` | 93.01 | 93.74 | 87.85 | | `hiresnet50.tv2_in1k` | 92.80 | 93.81 | 87.28 | | `hiresnet50.tv_in1k` | 92.71 | 94.09 | 86.16 | | `hivit_base_patch16_224.dino` | 88.00 | 91.10 | 86.47 | | `hivit_base_patch16_224.in1k_mocov3` | 90.19 | 93.07 | 87.73 | | `hivit_base_patch16_224.mae` | 83.62 | 92.79 | 85.26 | | `hivit_base_patch16_224.orig_in21k` | 89.32 | 91.73 | 90.97 | | `hivit_base_patch16_224_miil.in21k` | 90.49 | 92.13 | 90.21 | | `hivit_base_patch16_clip_224.laion2b` | 86.41 | 92.36 | 80.45 | | `hivit_base_patch16_siglip_224.v2_webli` | 93.40 | 95.06 | 87.81 | ## Requirements - `fgir-zoo` (`pip install git+https://github.com/arkel23/fgir-zoo.git`), which pins `timm==0.9.12` - `torch` (checked with 2.5.1) ## Citation ```bibtex @inproceedings{surya_revisiting_2025, title = {Revisiting the Backbone, Pretraining and Transferability for Hierarchical Fine-Grained Image Recognition}, author = {Surya, Augusto Christian and Rios, Edwin Arkel and Lai, Bo-Cheng and Hu, Min-Chun}, booktitle = {Computer Vision, Graphics, and Image Processing (CVGIP)}, year = {2025}, eprint = {TBA}, note = {A. C. Surya and E. A. Rios contributed equally. arXiv ID to be added.} } ```