Instructions to use MrFrogIsMe/vertag-explainer-lora-prompt-v0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use MrFrogIsMe/vertag-explainer-lora-prompt-v0 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
explainer LoRA, prompt v0 (VERTAG, ACCV 2026)
A LoRA adapter for Qwen/Qwen2.5-VL-7B-Instruct: given an applied trademark and the prior mark it was refused
over, it writes an examiner-style rationale (Traditional Chinese) for why the two marks are similar, item by item
over the visual likelihood-of-confusion factors. It is the grounded-explanation module of
VERTAG (ACCV 2026). The code, the exact prompt and both adapters are
described in explainer/.
The adapter of the paper's content rows (Tab. 4; Suppl. S7), including the region-evidence condition:
FADE's $C_{ij}$ correspondence text and matched region crops added to the two images. Region evidence does not
raise this fine-tuned model's element recall (0.584 → 0.559). For registration-number grounding use
MrFrogIsMe/vertag-explainer-lora.
Usage
git clone https://github.com/spaces-lalala/VERTAG.git
cd VERTAG/explainer
pip install -r requirements.txt # GPU with ~17 GB free (bf16)
hf download MrFrogIsMe/vertag-fade fade_dinov2_vitl14_reg.safetensors --local-dir ../FADE/checkpoints
python generate.py --applied applied.jpg --cited cited.jpg --condition C \
--fade-checkpoint ../FADE/checkpoints/fade_dinov2_vitl14_reg.safetensors \
--adapter MrFrogIsMe/vertag-explainer-lora-prompt-v0
--adapter takes this Hugging Face id directly; omit it to run the base model zero-shot.
Model card
- Source. The adapter files used in the paper, not retrained ones.
- Architecture. LoRA (r = 16, α = 32) on the q/k/v/o projections of the language model of
Qwen/Qwen2.5-VL-7B-Instruct; the vision tower is unchanged. Applied to the base model in bf16. - Training data. 6,000 applied → cited pairs from TIPO office actions published up to 2023, with the examiner's similarity paragraph as the target; QLoRA (4-bit), 2 epochs. The paper's evaluation cases are office actions after 2023.
- Limitations. Without the cited registration number (
--regno-evidence), a fine-tuned adapter almost always writes a fabricated one (paper Tab. 4: 100%). The output is automatically generated research text, not an examination opinion of TIPO and not legal advice. - Prompt. Trained with an earlier prompt that also asked about pronunciation; it is always run with the
current visual-only prompt, as in the paper. With images only, the two adapters reach almost the same element
recall (0.584 for this one, 0.581 for
vertag-explainer-lora).
| File | SHA256 |
|---|---|
adapter_model.safetensors |
3a64dfefc1896d851559a14fafd1f7e5e5958fea292832dee961e10549e3da50 |
License
The adapter is released under CC BY-NC 4.0; commercial use
is prohibited. It is applied to Qwen/Qwen2.5-VL-7B-Instruct (Apache-2.0, LICENSE-APACHE-2.0.txt), whose
terms also apply.
Citation
@inproceedings{yen2026vertag,
title = {{VERTAG}: Visual Examiner Rationales for Trademarks with Atomic Grounding --- A Confusion Benchmark, Faithful Retriever, and Explanation-Coverage Metric},
author = {Yen, Sheng-Yuan and Chou, Chia-Yi and Peng, Chi-Tse and Ye, Chian-Yu and Yu, Tsan-Wei and Ko, Chih-Chun and Wu, Yi-Chieh},
booktitle = {Proceedings of the Asian Conference on Computer Vision (ACCV)},
year = {2026}
}
- Downloads last month
- 20
Model tree for MrFrogIsMe/vertag-explainer-lora-prompt-v0
Base model
Qwen/Qwen2.5-VL-7B-Instruct