Papers
arxiv:2609.22323

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning

Published on Sep 16
· Submitted by
Neeraj Yadav
on Sep 23
Authors:

Abstract

Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator. Under a strictly matched, iso-episode-budget protocol (250 meta-training episodes, 5 canonical seeds, 600 evaluation episodes per seed), our architecture achieves 5-shot accuracy gains, consistent across all five seeds, over Prototypical Networks, Relation Networks, and MAML on both CIFAR-FS and MiniImageNet, while using 27-53% fewer parameters than any baseline. It also converges in fewer training episodes, generalizes better to an unseen fine-grained domain (CUB-200-2011 birds, zero retraining), and is more robust to 50% occlusion and 25% spatial translation than all three baselines. A series of falsification ablations - zeroing relational tokens at inference and retraining without them entirely - shows that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is. We report this honestly, together with a capacity sweep showing a genuine accuracy plateau near 22-35k parameters, and release full seed-level results and checkpoint hashes for reproducibility.

Community

Paper author Paper submitter

ALPINE is an ultra-lightweight few-shot image classification architecture designed for parameter- and sample-efficient learning. It combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator. Under a strictly matched 250-episode training budget, ALPINE achieves consistent 5-shot accuracy gains over Prototypical Networks, Relation Networks, and MAML on CIFAR-FS and MiniImageNet while using 27–53% fewer parameters. It also demonstrates improved robustness to occlusion and spatial translation and zero-retraining cross-domain generalization to CUB-200-2011. Ablation experiments further identify content-adaptive localization, rather than relational tokens, as the primary driver of the observed gains. Full checkpoints, seed-level results, and reproduction manifests are publicly released.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.22323 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.