TokenReductionPT / README.md
xedwin23x's picture
Add model card and manifest
c6e79ef verified
|
Raw
History Blame Contribute Delete
12.7 kB
metadata
pipeline_tag: image-classification
library_name: pytorch
tags:
  - fine-grained-image-recognition
  - image-classification
  - token-reduction
  - vision-transformer
  - timm

TokenReductionPT: token reduction and cross-layer cache aggregation for fine-grained recognition

These are the checkpoints behind Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image Recognition (arXiv:2501.00243), ICASSP 2025, and behind its extension, in preparation. Vision transformers are fine-tuned with a token-reduction method (EViT, TopK, DynamicViT, ATS, ToMe, PatchMerger, SiT, DPC-KNN and the FGIR-aware MAWS, DMAWS, EDAR, NFEDAR and GLSF criteria), with and without cross-layer cache aggregation (CLCA). CLCA adds a cross-layer aggregation head (CLA) and a cross-layer cache (CLC) that keeps information from discarded tokens. Code: arkel23/TokenReductionPT.

2656 checkpoints, one per configuration, each the last epoch of one training run. Each file is a torch.save dict with config (the full training configuration), model (the state dict), accuracy and epoch, with no optimizer state. File names are the runs' experiment-log names: dataset, model (reduction method as prefix), keep rate, cla for the aggregation head, clc_1 for the cache, and the serial. Load them with fgir-zoo. The collection groups this repo with the ICASSP 2025 paper.

Layout

One folder per serial. setting is bl (no CLCA), cla (aggregation head only) or clca (head and cache). "Extension" is the extension of the ICASSP 2025 paper, in preparation.

Folder What Used in Files Mean accuracy
serial_30 CLCA main results: EViT and the plain ViT-B, nine pretrainings, bl/cla/clca, 448 px ICASSP 2025, extension 405 58.04
serial_31 Plain ViT-B, two pretrainings, bl/cla, on CUB and Aircraft, 448 px ICASSP 2025, extension 8 89.74
serial_15 13 reduction methods x bl/clca, DeiT-T, 224 px extension 702 65.39
serial_16 13 reduction methods x bl/clca, DeiT-S, 224 px extension 702 71.68
serial_62 13 reduction methods x bl/clca, DeiT-T, 448 px extension 702 73.37
serial_408 TopK, DeiT-T, CUB, 288 px, keep rates 0.25-1.0 extension 4 77.01
serial_20 DMAWS with CLCA, six ViT-B pretrainings, cotton and soybean datasets, 448 px extension 93 75.47
serial_21 ImageNet-1k: EViT and the five FGIR-aware criteria, DeiT-T/S/B, 224 px extension 40 74.25

manifest.csv lists every file with its dataset, model, method, backbone, keep rate, setting, serial, seed, image size, class count, accuracy, the papers that use it, SHA-256 and size. The 224 px ablations of the ICASSP 2025 paper have no released checkpoints.

Load a checkpoint and classify an image

import torch
from PIL import Image
from torchvision import transforms
from fgir_zoo import tokenred

model = tokenred.create_model('serial_62/cub_evit_deit_tiny_patch16_224.fb_in1k_0.7_cla_clc_1_62')
cfg = model.config
# evaluation at 448 px resizes to a square without cropping
tf = transforms.Compose([
    transforms.Resize((cfg.input_size, cfg.input_size),
                      interpolation=transforms.InterpolationMode.BICUBIC),
    transforms.ToTensor(),
    transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
# a CUB-200-2011 test image, class index 50 (051.Horned_Grebe)
x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0)
with torch.no_grad():
    logits = model(x)
print(logits.argmax(-1).item(), logits.softmax(-1).max().item())  # 50 0.9315

Accuracy of the released checkpoints

Top-1 accuracy (%) is the run's own test accuracy after the last epoch, as logged in training. The papers average three seeds; where a configuration ran with several seeds, the released file is the seed with the highest accuracy, so the papers' tables can differ slightly. Per-file accuracy is in manifest.csv; 1 file has no recorded accuracy (serial_16/cub_edar_deit_small_patch16_224.fb_in1k_0.7_16). The table below covers serial_30, the ICASSP 2025 main results; method none is the plain ViT without token reduction.

Method, keep rate, setting Backbone cotton soyageing soygene soyglobal soylocal
evit 0.1 bl deit3_base_patch16_224.fb_in1k 18.33 66.48 57.99 30.00 29.17
evit 0.1 bl deit3_base_patch16_224.fb_in22k_ft_in1k 20.83 64.36 38.25 12.13 21.00
evit 0.1 bl deit_base_patch16_224.fb_in1k 22.08 71.70 56.81 16.94 33.67
evit 0.1 bl vit_base_patch16_224.dino 27.08 72.57 47.86 29.72 29.00
evit 0.1 bl vit_base_patch16_224.in1k_mocov3 28.75 66.10 40.89 3.85 29.50
evit 0.1 bl vit_base_patch16_224.mae 10.42 22.59 37.31 3.89 18.00
evit 0.1 bl vit_base_patch16_224.orig_in21k 35.42 73.29 57.10 39.03 33.67
evit 0.1 bl vit_base_patch16_224_miil.in21k 17.50 69.52 49.02 21.57 27.67
evit 0.1 bl vit_base_patch16_clip_224.laion2b 16.67 58.95 22.42 3.90 26.50
evit 0.1 cla deit3_base_patch16_224.fb_in1k 46.25 86.36 76.15 60.27 51.67
evit 0.1 cla deit3_base_patch16_224.fb_in22k_ft_in1k 51.67 87.21 77.17 51.48 50.50
evit 0.1 cla deit_base_patch16_224.fb_in1k 40.42 83.15 71.96 46.94 48.33
evit 0.1 cla vit_base_patch16_224.dino 52.08 85.23 70.52 58.57 47.33
evit 0.1 cla vit_base_patch16_224.in1k_mocov3 40.00 81.03 66.44 41.76 44.17
evit 0.1 cla vit_base_patch16_224.mae 43.33 78.73 66.72 58.00 17.83
evit 0.1 cla vit_base_patch16_224.orig_in21k 55.83 85.45 70.82 41.73 51.67
evit 0.1 cla vit_base_patch16_224_miil.in21k 46.25 79.88 64.35 52.92 52.50
evit 0.1 cla vit_base_patch16_clip_224.laion2b 55.00 87.68 72.70 65.62 59.67
evit 0.1 clca deit3_base_patch16_224.fb_in1k 61.25 87.31 69.35 61.99 57.50
evit 0.1 clca deit3_base_patch16_224.fb_in22k_ft_in1k 56.67 88.42 76.72 62.56 53.67
evit 0.1 clca deit_base_patch16_224.fb_in1k 53.33 85.60 74.65 59.03 56.83
evit 0.1 clca vit_base_patch16_224.dino 54.17 85.49 71.61 60.80 48.67
evit 0.1 clca vit_base_patch16_224.in1k_mocov3 49.17 82.89 69.15 49.48 47.50
evit 0.1 clca vit_base_patch16_224.mae 25.42 85.54 65.34 24.25 42.67
evit 0.1 clca vit_base_patch16_224.orig_in21k 54.58 83.78 71.34 48.74 51.67
evit 0.1 clca vit_base_patch16_224_miil.in21k 55.00 85.07 71.95 55.71 51.17
evit 0.1 clca vit_base_patch16_clip_224.laion2b 54.58 89.35 76.97 58.10 62.17
evit 0.7 bl deit3_base_patch16_224.fb_in1k 37.50 76.81 66.72 29.38 33.83
evit 0.7 bl deit3_base_patch16_224.fb_in22k_ft_in1k 35.00 73.41 53.06 6.60 32.67
evit 0.7 bl deit_base_patch16_224.fb_in1k 54.58 82.26 69.75 33.35 47.83
evit 0.7 bl vit_base_patch16_224.dino 40.42 76.95 55.75 33.20 41.00
evit 0.7 bl vit_base_patch16_224.in1k_mocov3 55.83 76.22 58.58 2.10 46.33
evit 0.7 bl vit_base_patch16_224.mae 15.42 73.41 20.88 21.93 36.33
evit 0.7 bl vit_base_patch16_224.orig_in21k 59.17 79.64 60.28 39.54 46.33
evit 0.7 bl vit_base_patch16_224_miil.in21k 40.00 76.32 48.94 32.01 37.33
evit 0.7 bl vit_base_patch16_clip_224.laion2b 7.50 51.37 15.25 8.89 9.67
evit 0.7 cla deit3_base_patch16_224.fb_in1k 62.08 88.77 80.97 58.51 55.33
evit 0.7 cla deit3_base_patch16_224.fb_in22k_ft_in1k 61.67 86.89 76.80 53.58 53.83
evit 0.7 cla deit_base_patch16_224.fb_in1k 67.92 87.82 79.40 59.15 58.83
evit 0.7 cla vit_base_patch16_224.dino 68.33 88.40 73.26 60.75 57.83
evit 0.7 cla vit_base_patch16_224.in1k_mocov3 65.83 86.83 74.32 53.13 58.33
evit 0.7 cla vit_base_patch16_224.mae 59.58 90.67 77.47 51.38 57.83
evit 0.7 cla vit_base_patch16_224.orig_in21k 58.75 86.51 71.75 55.26 55.17
evit 0.7 cla vit_base_patch16_224_miil.in21k 59.58 87.09 73.39 64.00 59.00
evit 0.7 cla vit_base_patch16_clip_224.laion2b 60.83 89.96 40.57 57.24 61.00
evit 0.7 clca deit3_base_patch16_224.fb_in1k 66.67 88.08 80.18 59.86 55.83
evit 0.7 clca deit3_base_patch16_224.fb_in22k_ft_in1k 62.50 86.51 73.76 52.05 59.00
evit 0.7 clca deit_base_patch16_224.fb_in1k 65.83 88.65 78.93 62.37 60.17
evit 0.7 clca vit_base_patch16_224.dino 67.08 88.14 75.12 64.65 59.17
evit 0.7 clca vit_base_patch16_224.in1k_mocov3 65.83 87.15 75.26 55.90 59.00
evit 0.7 clca vit_base_patch16_224.mae 57.92 90.51 74.34 41.40 54.83
evit 0.7 clca vit_base_patch16_224.orig_in21k 64.17 85.84 74.65 56.81 57.33
evit 0.7 clca vit_base_patch16_224_miil.in21k 65.00 86.59 72.46 60.17 58.83
evit 0.7 clca vit_base_patch16_clip_224.laion2b 68.33 91.27 79.48 62.64 62.83
none bl deit3_base_patch16_224.fb_in1k 47.92 76.40 66.23 26.32 33.83
none bl deit3_base_patch16_224.fb_in22k_ft_in1k 35.00 72.59 52.88 15.38 32.33
none bl deit_base_patch16_224.fb_in1k 56.25 81.90 72.84 39.34 44.00
none bl vit_base_patch16_224.dino 43.75 75.56 55.42 35.72 39.50
none bl vit_base_patch16_224.in1k_mocov3 55.42 76.71 56.56 1.86 46.00
none bl vit_base_patch16_224.mae 28.75 74.12 30.82 10.56 22.50
none bl vit_base_patch16_224.orig_in21k 58.75 79.70 61.90 37.39 49.00
none bl vit_base_patch16_224_miil.in21k 31.67 75.21 46.06 35.31 36.50
none bl vit_base_patch16_clip_224.laion2b 7.08 40.95 12.69 9.61 14.17
none cla deit3_base_patch16_224.fb_in1k 66.67 87.74 80.88 58.82 57.00
none cla deit3_base_patch16_224.fb_in22k_ft_in1k 62.08 85.58 75.52 52.22 51.67
none cla deit_base_patch16_224.fb_in1k 66.25 87.49 80.36 63.83 57.83
none cla vit_base_patch16_224.dino 68.33 88.02 75.62 61.30 57.83
none cla vit_base_patch16_224.in1k_mocov3 67.50 86.69 76.13 55.35 59.33
none cla vit_base_patch16_224.mae 65.00 90.04 77.23 51.41 60.50
none cla vit_base_patch16_224.orig_in21k 65.00 87.17 72.01 56.66 57.83
none cla vit_base_patch16_224_miil.in21k 60.42 87.19 74.59 64.52 56.17
none cla vit_base_patch16_clip_224.laion2b 53.33 90.93 76.59 55.21 59.33
none clca deit3_base_patch16_224.fb_in1k 67.50 88.04 80.71 61.75 59.50
none clca deit3_base_patch16_224.fb_in22k_ft_in1k 59.58 87.21 79.00 59.43 50.33
none clca deit_base_patch16_224.fb_in1k 69.58 88.79 81.04 62.49 63.17
none clca vit_base_patch16_224.dino 68.33 88.08 72.49 64.74 59.50
none clca vit_base_patch16_224.in1k_mocov3 65.00 87.37 76.76 57.29 58.83
none clca vit_base_patch16_224.mae 52.92 90.32 76.42 53.32 60.50
none clca vit_base_patch16_224.orig_in21k 65.00 86.26 74.55 58.46 57.00
none clca vit_base_patch16_224_miil.in21k 61.67 87.17 74.73 61.30 56.83
none clca vit_base_patch16_clip_224.laion2b 65.83 90.71 77.13 62.95 63.50

Requirements

  • fgir-zoo (pip install git+https://github.com/arkel23/fgir-zoo.git), which pins timm==0.9.12
  • torch (checked with 2.5.1)

Citation

@inproceedings{rios_cross-layer_2025,
  title         = {Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image Recognition},
  author        = {Rios, Edwin Arkel and Yuanda, Jansen Christopher and Ghanz, Vincent Leon and
                   Yu, Cheng-Wei and Lai, Bo-Cheng and Hu, Min-Chun},
  booktitle     = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year          = {2025},
  eprint        = {2501.00243},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2501.00243},
  url           = {https://arxiv.org/abs/2501.00243}
}