docverify β€” payslip tampering detection

A forensic segmentation network that finds where a payslip was edited β€” copy-move, splice from another document, or a re-typed amount β€” and the small models and thresholds that turn its pixel map into a verdict an analyst can act on.

Part of docverify, a complete system with arithmetic checks, PDF forensics and a local LLM note: project Β· YOLO field detector

A splice on a phone photo, found

Files

File What it is
tampernet.onnx the served model β€” ONNX FP32, 1.56Γ— faster than PyTorch on CPU, identical decisions
tampernet.pt the same weights for PyTorch
threshold.json page threshold 7.86 (5% false alarms on genuine payslips), field threshold 9.19 (1% of genuine pages rated high risk)
siglip_probe.json linear probe on frozen Google SigLIP image embeddings: payslip / RIB / invoice
version.txt release 2

Model

TamperNet (~33M parameters): a pretrained ConvNeXt encoder on RGB, plus two forensic inputs β€” fixed SRM noise residual filters (edits break the sensor/print noise) and a JPEG/DCT branch (a pasted region carries its own 8Γ—8 compression grid) β€” fused into a U-Net-style decoder that outputs one logit per pixel. The page score is the mean of the top-200 pixel logits.

Inference runs at full resolution (1000 Γ— 1414 px pages, padded to a multiple of 32 to keep the JPEG grid).

Results

Frozen test split: 793 synthetic payslips (390 genuine, 403 forged), never seen in training.

Metric v2 (this release) v1 (U-Net)
ROC-AUC 0.902 0.877
Forgeries caught at 1% false alarms 61% 55%
Forgeries caught at 5% false alarms 72% 66%
Pixel F1 on forged pages 0.81 0.77
Pixel IoU 0.68 0.62
Segment (recall at 5% false alarms) v2
Splice (pasted from another document) 90%
Copy-move (inside the document) 77%
Rewritten amount 50%
Scanned 77%
Phone photo 63%

v1 vs v2 ROC and scores Pixel localisation

Rewritten amounts are the hard case for pixels alone. In the full system, seven payslip arithmetic rules raise recall on them from 56% to 92%, and on all forgeries to 95% at 1.7% false alarms (240-document OCR subset):

Image vs image + arithmetic

Speed

A full page needs 4 GB of RAM and ~5 s on 4 CPU threads (0.44 s on a T4).

Throughput

Limits

  • Synthetic data only. Trained and evaluated on generated French payslips (Faker fr_FR, fake people and companies). Never measured on real documents.
  • Built for one layout family (French payslips). Other documents need adaptation; the forensic branches are generic.
  • It flags; it does not decide. In docverify, a human analyst always makes the decision.
  • EU AI Act: fraud detection is excluded from the Annex III 5(b) high-risk category. Using this score directly to grant or refuse credit would not be.

License

Apache-2.0. The ConvNeXt encoder was initialised from timm ImageNet weights (Apache-2.0). The SigLIP probe applies to Google's SigLIP (Apache-2.0), downloaded from Google's repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support