solvi extract-receipts

solvi-ai/extract-base fine-tuned on receipts — an example of adapting the general field-by-description extractor to one document type. Use it with solvi: each field comes back as the exact span of the OCR text (or "absent"), and plain Python rules decide.

from solvi.extract_long import LongSpanExtractor
ex = LongSpanExtractor.load("solvi-ai/extract-receipts")      # pip install "solvi[model]"
total = ex.field("total", "the total amount to pay")         # an @extract part for a solvi Catalog
print(total(receipt_text))                                    # Quote(value, start, end, confidence)

Per-field thresholds for total, subtotal, tax, cash and change (with the descriptions below) are stored in solvi_extract.json: "the total amount to pay", "the subtotal before tax and service", "the tax amount", "the cash amount given by the customer", "the change returned to the customer".

Training

Two epochs on the 800 training receipts of CORD v2 (CC BY 4.0): all 22 field types with written descriptions (items, prices, quantities, subtotal, service, tax, discount, total, cash, change, card and e-money payments, item counts) plus the five fields above. No SROIE data was used.

Evaluation (one seed)

test result
CORD test (100 receipts): typed questions via solvi rules — total band, tax charged, paid in cash, change correct 97.3%; ECE 0.011; 98.6% of questions answered at ≥ 99% precision
SROIE test (361 receipts, never trained): company / date / total by description 34% / 87% / 68%

It is strong on amounts (total, tax, cash, change) and dates. It lost the base model's zero-shot skill on shop names (company 71% → 34%) because CORD does not label shop names — use solvi-ai/extract-base or fine-tune on your own receipts (~25–100 labeled) for fields CORD lacks.

Limitations

English/Indonesian-style receipts from OCR text; layouts far from CORD are untested. Confidences are span scores — calibrate per question with System.calibrate.

ONNX (browser and CPU)

onnx/model_fp16.onnx — the encoder and span head in one graph (inputs input_ids, attention_mask int64 [batch, length] → logits float [batch, length, 2]: start and end scores), fp16 weights with float32 inputs and outputs. On 175 test fields it returns exactly the same spans as the PyTorch model. Runs with onnxruntime (CPU) and onnxruntime-web (WebGPU); windowing and span decoding follow LongSpanExtractor.predict. Export your own fine-tuned model with tools/export_onnx.py in the solvi repo. (Dynamic int8 quantization is not provided: it changed half of the spans.)

Downloads last month
32
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for solvi-ai/extract-receipts

Quantized
(1)
this model

Dataset used to train solvi-ai/extract-receipts

Spaces using solvi-ai/extract-receipts 2