Instructions to use Nafety/off-ingredient-linker with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Nafety/off-ingredient-linker with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") model = PeftModel.from_pretrained(base_model, "Nafety/off-ingredient-linker") - Notebooks
- Google Colab
- Kaggle
off-ingredient-linker
A LoRA adapter for Qwen2.5-1.5B-Instruct. It proposes a fix for ingredient names that the Open Food Facts parser does not recognize.
Code, training data pipeline and evaluation: https://github.com/Nafety/ingredient-linker
What it does
The model receives an unrecognized name, the declared language of the ingredient list, the surrounding list, and the 10 closest taxonomy entries found by a multilingual embedding model (intfloat/multilingual-e5-small). It answers with one JSON decision:
| Decision | Example |
|---|---|
{"type": "existing", "id": ...}: an existing entry (other language, typo, OCR error) |
paprikapulver → en:paprika-powder |
{"type": "new", "parent": ...}: missing from the taxonomy, with its closest parent |
triple filtered purified water → parent en:filtered-water |
{"type": "not_ingredient"}: nutrition facts, addresses, fragments |
energy kj energy kcal fat |
{"type": "structure"}: list phrases |
tocopherols added to maintain freshness |
The prompt format matters: build it with the code from the GitHub repository rather than by hand.
git clone https://github.com/Nafety/ingredient-linker && cd ingredient-linker
pip install -r requirements.txt
python infer.py "glyserol" "paprikapulver" --adapter Nafety/off-ingredient-linker
Training
- Data: 37,214 examples built automatically from the Open Food Facts export and taxonomies:
- links made by the OFF parser;
- taxonomy names damaged on purpose;
- taxonomy leaves hidden to simulate new ingredients;
- rule-based noise and list phrases.
- Method: QLoRA. The base model is loaded in 4-bit NF4, and LoRA (rank 16, alpha 32) is applied to every linear layer: 18.5 M trainable parameters (1.2 %).
- Run: 20,000 examples, 1 epoch, learning rate 2e-4, loss on the answer only. It took about 3 h on an 8 GB RTX 4060 Laptop.
Evaluation
The test set has 300 real unrecognized names, each found in at least 10 products, annotated by hand (293 scored). None of them is in the training data. Strict means the right decision and the right entry (or parent).
| Approach | Right decision | Strict | Allergen recall | False allergen alerts |
|---|---|---|---|---|
| Current OFF parser (these are its failures) | 0 % | 0 % | 0 % | 0 |
| Best non-LLM baseline (exact > fuzzy > retrieval) | 70.6 % | 62.5 % | 68.6 % | 11 |
| Qwen2.5-1.5B-Instruct, not fine-tuned | 37.5 % | 33.1 % | 40.0 % | 1 |
| This adapter | 78.2 % | 68.6 % | 91.4 % | 4 |
Limitations
- New ingredients missing from the taxonomy are rarely handled well (12.5 % strict, on 16 cases).
- Small test set: 293 cases, so differences of a few points are not significant.
- Use as suggestions only. The output should be reviewed by a contributor before any change to the taxonomy or to products.
- Downloads last month
- 1