Instructions to use berkerdooo/gemma-4-e2b-receipt-json with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use berkerdooo/gemma-4-e2b-receipt-json with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-E2B-it") model = PeftModel.from_pretrained(base_model, "berkerdooo/gemma-4-e2b-receipt-json") - Notebooks
- Google Colab
- Kaggle
Gemma 4 receipt images to JSON
A LoRA adapter for Gemma 4 E2B that extracts Indonesian receipt images into JSON: item names, quantities and prices, subtotal, tax and total. Printed amounts remain strings, preserving the source formatting. The repository contains the adapter and processor; loading it also requires the base model.
Evaluation against the base model
| Metric | Base | Fine tuned | Change |
|---|---|---|---|
| Strict field F1 | 0.00% | 84.21% | +84.21 points |
| Content field F1 (fences removed) | 51.59% | 84.21% | +32.62 points |
| Total exact match | 0.00% | 86.00% | +86.00 points |
| Valid JSON rate | 0.00% | 100.00% | +100.00 points |
Greedy generation on all 100 held-out receipt images, using the same prompt, up to 280 image soft tokens and a 1024-token output limit for both models. Field F1 compares exact JSON path and serialized value pairs, including item indices and null fields. Total exact match compares the total string. Invalid JSON receives an empty prediction in the strict scores. Content field F1 is a separate diagnostic that removes a single Markdown code fence before parsing; it makes no repairs to field values. The base model commonly emits fenced JSON, so strict scoring understates its extraction ability. Training uses receipt pixels as input; target annotations are never included in the prompt. Validation loss on 40 receipts selects the saved epoch.
These results describe the fixed held-out subset from the public dataset. They do not measure general performance across other domains. evaluation.json and both prediction files provide the underlying evidence.
Training and data
Base: google/gemma-4-E2B-it. Dataset: naver-clova-ix/cord-v2.
Training used 798 examples, seed 42 and 200 optimizer steps on one NVIDIA L40S. The three task jobs ran concurrently. The initial training job took 36.9 minutes, including model loading, evaluation and time waiting for the other base evaluations. Subsequent evaluation and reload checks are recorded separately. This is not an isolated training-speed benchmark. Training used PyTorch 2.9.1 with CUDA 12.8, Transformers 5.16.0 and the package pins included in the repository. Arguments and environment details are saved alongside the weights.
{
"learning_rate": 0.0002,
"per_device_train_batch_size": 1,
"gradient_accumulation_steps": 8,
"num_train_epochs": 2,
"lr_scheduler_type": "cosine",
"bf16": true,
"seed": 42,
"actual_training_rows": 798,
"optimizer_steps": 200,
"actual_epochs": 2.0,
"base_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"dataset_revision": "7f0115a4b758a71d6473b8d085751692da2fef98",
"selected_checkpoint": "checkpoint-200",
"lora_rank": 16,
"lora_alpha": 32,
"lora_dropout": 0.05,
"image_soft_tokens": 280,
"frozen_components": "All base weights; adapters train in language decoder projections only."
}
Dataset split counts: {"test": 100, "validation": 100, "train": 798}. All official splits; exact duplicate image pixels removed. Indonesian shop/restaurant receipts. Receipt images supply input; gt_parse annotations supply assistant targets and evaluation references.
Usage
Install the pinned packages in reproduce/requirements-gpu.txt with a compatible PyTorch build, then run:
import torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForMultimodalLM
from peft import PeftModel
repo = "berkerdooo/gemma-4-e2b-receipt-json"
processor = AutoProcessor.from_pretrained(repo)
base = AutoModelForMultimodalLM.from_pretrained(
"google/gemma-4-E2B-it", revision="3e22461f65e89153144f8adb70e3b8c2cc9845a7",
dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(base, repo)
prompt = "Extract this receipt as JSON with exactly these keys: items (array of objects with name, quantity, price), subtotal, tax, total. Preserve the printed number strings without converting currency. Use null for missing fields. Return only JSON."
messages = [{"role":"user", "content":[
{"type":"image", "image":Image.open("receipt.jpg").convert("RGB")},
{"type":"text", "text":prompt}
]}]
inputs = processor.apply_chat_template(messages, tokenize=True, add_generation_prompt=True,
return_dict=True, return_tensors="pt", enable_thinking=False,
processor_kwargs={"images_kwargs":{"max_soft_tokens":280}}).to(model.device)
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=1024, do_sample=False)
print(processor.tokenizer.decode(output[0, inputs.input_ids.shape[1]:], skip_special_tokens=True))
Scope and limitations
CORD contains Indonesian shop and restaurant receipts. Results do not establish accuracy on Turkish or English invoices, other currencies, handwriting or poor scans. The official test contains 100 receipts. Item order affects the field metric; outputs are capped at 1024 tokens. Missing annotation fields become null.
Reproduce the run
The reproduce folder contains the training script, shared helpers, preparation code, package pins and revision manifest. Run PREPARE_TASKS=receipts python prepare.py to prepare this task locally, then copy the prepared data and any downloaded base weights into the GPU environment. The scripts compare the unchanged base with the fine-tuned model and coordinate concurrent starts through base_ready markers. For a standalone run, set FINETUNE_TASKS=receipts before executing the training script. Training and evaluation use the official splits with the documented fixed-seed sampling.
After training, run python analyze_predictions.py outputs/receipts to compute the separate content metric.
License and attribution
Weights and adapters use Apache License 2.0.
Training data: CORD v2 by NAVER CLOVA, under CC BY 4.0. Receipt annotations were transformed into the schema above; JPEG images were prepared for training. Dataset citation: Park et al., “CORD: A Consolidated Receipt Dataset for Post-OCR Parsing,” 2019.
Fine-tune prepared for berkerdooo. Base model credit belongs to Google DeepMind and the Gemma team.
- Downloads last month
- 6