off-ingredient-linker

A LoRA adapter for Qwen2.5-1.5B-Instruct. It proposes a fix for ingredient names that the Open Food Facts parser does not recognize.

Code, training data pipeline and evaluation: https://github.com/Nafety/ingredient-linker

What it does

The model receives an unrecognized name, the declared language of the ingredient list, the surrounding list, and the 10 closest taxonomy entries found by a multilingual embedding model (intfloat/multilingual-e5-small). It answers with one JSON decision:

Decision Example
{"type": "existing", "id": ...}: an existing entry (other language, typo, OCR error) paprikapulver → en:paprika-powder
{"type": "new", "parent": ...}: missing from the taxonomy, with its closest parent triple filtered purified water → parent en:filtered-water
{"type": "not_ingredient"}: nutrition facts, addresses, fragments energy kj energy kcal fat
{"type": "structure"}: list phrases tocopherols added to maintain freshness

The prompt format matters: build it with the code from the GitHub repository rather than by hand.

git clone https://github.com/Nafety/ingredient-linker && cd ingredient-linker
pip install -r requirements.txt
python infer.py "glyserol" "paprikapulver" --adapter Nafety/off-ingredient-linker

Training

  • Data: 37,214 examples built automatically from the Open Food Facts export and taxonomies:
    • links made by the OFF parser;
    • taxonomy names damaged on purpose;
    • taxonomy leaves hidden to simulate new ingredients;
    • rule-based noise and list phrases.
  • Method: QLoRA. The base model is loaded in 4-bit NF4, and LoRA (rank 16, alpha 32) is applied to every linear layer: 18.5 M trainable parameters (1.2 %).
  • Run: 20,000 examples, 1 epoch, learning rate 2e-4, loss on the answer only. It took about 3 h on an 8 GB RTX 4060 Laptop.

Evaluation

The test set has 300 real unrecognized names, each found in at least 10 products, annotated by hand (293 scored). None of them is in the training data. Strict means the right decision and the right entry (or parent).

Approach Right decision Strict Allergen recall False allergen alerts
Current OFF parser (these are its failures) 0 % 0 % 0 % 0
Best non-LLM baseline (exact > fuzzy > retrieval) 70.6 % 62.5 % 68.6 % 11
Qwen2.5-1.5B-Instruct, not fine-tuned 37.5 % 33.1 % 40.0 % 1
This adapter 78.2 % 68.6 % 91.4 % 4

Limitations

  • New ingredients missing from the taxonomy are rarely handled well (12.5 % strict, on 16 cases).
  • Small test set: 293 cases, so differences of a few points are not significant.
  • Use as suggestions only. The output should be reviewed by a contributor before any change to the taxonomy or to products.
Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nafety/off-ingredient-linker

Adapter
(1521)
this model

Dataset used to train Nafety/off-ingredient-linker