solvi extract-receipts
solvi-ai/extract-base fine-tuned on receipts — an example of adapting the general field-by-description extractor to one document type. Use it with solvi: each field comes back as the exact span of the OCR text (or "absent"), and plain Python rules decide.
from solvi.extract_long import LongSpanExtractor
ex = LongSpanExtractor.load("solvi-ai/extract-receipts") # pip install "solvi[model]"
total = ex.field("total", "the total amount to pay") # an @extract part for a solvi Catalog
print(total(receipt_text)) # Quote(value, start, end, confidence)
Per-field thresholds for total, subtotal, tax, cash and change (with the descriptions below) are stored in
solvi_extract.json: "the total amount to pay", "the subtotal before tax and service", "the tax amount", "the cash amount given
by the customer", "the change returned to the customer".
Training
Two epochs on the 800 training receipts of CORD v2 (CC BY 4.0): all 22 field types with written descriptions (items, prices, quantities, subtotal, service, tax, discount, total, cash, change, card and e-money payments, item counts) plus the five fields above. No SROIE data was used.
Evaluation (one seed)
| test | result |
|---|---|
| CORD test (100 receipts): typed questions via solvi rules — total band, tax charged, paid in cash, change correct | 97.3%; ECE 0.011; 98.6% of questions answered at ≥ 99% precision |
| SROIE test (361 receipts, never trained): company / date / total by description | 34% / 87% / 68% |
It is strong on amounts (total, tax, cash, change) and dates. It lost the base model's zero-shot skill on shop names
(company 71% → 34%) because CORD does not label shop names — use solvi-ai/extract-base or fine-tune on your own receipts
(~25–100 labeled) for fields CORD lacks.
Limitations
English/Indonesian-style receipts from OCR text; layouts far from CORD are untested. Confidences are span scores — calibrate
per question with System.calibrate.
ONNX (browser and CPU)
onnx/model_fp16.onnx — the encoder and span head in one graph (inputs input_ids, attention_mask int64 [batch, length] →
logits float [batch, length, 2]: start and end scores), fp16 weights with float32 inputs and outputs. On 175 test fields it
returns exactly the same spans as the PyTorch model. Runs with onnxruntime (CPU) and onnxruntime-web (WebGPU); windowing and
span decoding follow LongSpanExtractor.predict. Export your own fine-tuned model with tools/export_onnx.py in the solvi repo.
(Dynamic int8 quantization is not provided: it changed half of the spans.)
- Downloads last month
- 32