Instructions to use MrFrogIsMe/vertag-explainer-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use MrFrogIsMe/vertag-explainer-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 3,123 Bytes
9a1efe9 3767634 9a1efe9 9dc8e7c 9a1efe9 3767634 9a1efe9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 | ---
license: cc-by-nc-4.0
base_model: Qwen/Qwen2.5-VL-7B-Instruct
library_name: peft
pipeline_tag: image-text-to-text
language:
- zh
tags:
- lora
- peft
- trademark
- legal
- vertag
---
# explainer LoRA (VERTAG, ACCV 2026)
A LoRA adapter for `Qwen/Qwen2.5-VL-7B-Instruct`: given an applied trademark and the prior mark it was refused
over, it writes an examiner-style rationale (Traditional Chinese) for why the two marks are similar, item by item
over the visual likelihood-of-confusion factors. It is the grounded-explanation module of
[VERTAG](https://github.com/spaces-lalala/VERTAG) (ACCV 2026). The code, the exact prompt and both adapters are
described in [`explainer/`](https://github.com/spaces-lalala/VERTAG/tree/main/explainer).
The **default** adapter. It was trained with the visual-only prompt that `generate.py` runs, and it is the
adapter of the paper's registration-number grounding (Tab. 4, registration-number row: fabricated registration
numbers fall from 100% to 5% at no cost in content).
## Usage
```bash
git clone https://github.com/spaces-lalala/VERTAG.git
cd VERTAG/explainer
pip install -r requirements.txt # GPU with ~17 GB free (bf16)
python generate.py --applied applied.jpg --cited cited.jpg --regno 00820591 \
--condition C --regno-evidence --adapter MrFrogIsMe/vertag-explainer-lora
```
`--adapter` takes this Hugging Face id directly; omit it to run the base model zero-shot.
## Model card
- **Source.** The adapter files used in the paper, not retrained ones.
- **Architecture.** LoRA (r = 16, α = 32) on the q/k/v/o projections of the language model of
`Qwen/Qwen2.5-VL-7B-Instruct`; the vision tower is unchanged. Applied to the base model in bf16.
- **Training data.** 6,000 applied → cited pairs from TIPO office actions published up to 2023, with the
examiner's similarity paragraph as the target; QLoRA (4-bit), 2 epochs. The paper's evaluation cases are
office actions after 2023.
- **Limitations.** Without the cited registration number (`--regno-evidence`), a fine-tuned adapter almost
always writes a fabricated one (paper Tab. 4: 100%). The output is automatically generated research text,
not an examination opinion of TIPO and not legal advice.
- **Prompt.** Trained with the prompt it is run with.
| File | SHA256 |
|---|---|
| `adapter_model.safetensors` | `43c8d681f3f822303c885c5b8080a2a54098f5b8d8868ea518d1b6c969ea6597` |
## License
The adapter is released under [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/); commercial use
is prohibited. It is applied to `Qwen/Qwen2.5-VL-7B-Instruct` (Apache-2.0, `LICENSE-APACHE-2.0.txt`), whose
terms also apply.
## Citation
```bibtex
@inproceedings{yen2026vertag,
title = {{VERTAG}: Visual Examiner Rationales for Trademarks with Atomic Grounding --- A Confusion Benchmark, Faithful Retriever, and Explanation-Coverage Metric},
author = {Yen, Sheng-Yuan and Chou, Chia-Yi and Peng, Chi-Tse and Ye, Chian-Yu and Yu, Tsan-Wei and Ko, Chih-Chun and Wu, Yi-Chieh},
booktitle = {Proceedings of the Asian Conference on Computer Vision (ACCV)},
year = {2026}
}
```
|