FADE checkpoints (VERTAG, ACCV 2026)
FADE (Faithful Additive DEscriptor) is the examiner-confusion retriever of VERTAG (ACCV 2026). A light additive attention-pooling head on a mostly frozen backbone pools an image's patch tokens $v_i$ into $d=\sum_i a_i v_i$, so the cosine between two trademarks decomposes exactly into patch-pair contributions $C_{ij}$ that sum to the retrieval score itself.
| File | Backbone | Use it for | G1 R@100 |
|---|---|---|---|
fade_siglip_so400m.safetensors |
SigLIP-SO400M/14, 224 px | retrieval: the strongest retriever in the paper (Tab. 3) | 0.609 |
fade_dinov2_vitl14_reg.safetensors |
DINOv2-L/14 with registers, 224 px | the model analysed in the paper (Tab. 2, Fig. 3) and the source of the explainer's region evidence | 0.146 |
Usage
The checkpoints are loaded by the code in the VERTAG repository.
git clone https://github.com/spaces-lalala/VERTAG.git
cd VERTAG/FADE
pip install -r requirements.txt
hf download MrFrogIsMe/vertag-fade --include "*.safetensors" --local-dir checkpoints
python explain.py --checkpoint checkpoints/fade_dinov2_vitl14_reg.safetensors \
--applied applied.jpg --cited cited.jpg
python retrieve.py --checkpoint checkpoints/fade_siglip_so400m.safetensors \
--gallery /path/to/gallery_dir --query query.jpg --explain 3
A file stores only the tensors training changed; the rest of the backbone is downloaded from its official
pretrained weights on first use (DINOv2 from torch.hub, SigLIP from Hugging Face).
Model card
- Source. These are the paper's models, not retrained ones. Each file is a slim export of the paper's checkpoint, and every tensor was checked against the original checkpoint when it was exported.
- Architecture. Only the last two transformer blocks, the backbone's final norm and the FADE head were trained.
- Training data. METU-v2 copy-detection pairs (its 417 queries excluded) mixed with examiner pairs (applied mark → cited mark) from TIPO office actions published up to 2023. The benchmark queries are the office actions after 2023. Evaluation on METU-v2 is therefore not zero-shot. Some prior marks cited against benchmark queries also occur in the training pairs; the paper reports the effect (Suppl. S1).
- Intended use. Research on trademark retrieval and its explanation. A similarity score or a $C_{ij}$ map is not a legal assessment of likelihood of confusion.
| File | SHA256 |
|---|---|
fade_dinov2_vitl14_reg.safetensors |
41702efeaa7ec5dc2350173996aefd6ef2805ef063e2af45f72417dcc3eadce6 |
fade_siglip_so400m.safetensors |
d03b6aa280503a20947d8762dca2a1dbd946a54bdd2eef2c9515f0881e0a1a05 |
License
The weights are released under CC BY-NC 4.0;
commercial use is prohibited. They are derived from DINOv2 and SigLIP, both Apache-2.0
(LICENSE-APACHE-2.0.txt).
Citation
@inproceedings{yen2026vertag,
title = {{VERTAG}: Visual Examiner Rationales for Trademarks with Atomic Grounding --- A Confusion Benchmark, Faithful Retriever, and Explanation-Coverage Metric},
author = {Yen, Sheng-Yuan and Chou, Chia-Yi and Peng, Chi-Tse and Ye, Chian-Yu and Yu, Tsan-Wei and Ko, Chih-Chun and Wu, Yi-Chieh},
booktitle = {Proceedings of the Asian Conference on Computer Vision (ACCV)},
year = {2026}
}
Model tree for MrFrogIsMe/vertag-fade
Base model
facebook/dinov2-with-registers-large