Model description
Model Name: multicentury-ppocrv6-medium
Model Version: PP-OCRv6_medium_rec_supermalli_widerv4_nocyrillic
Model Type: PP-OCRv6
Base Model: PaddlePaddle/PP-OCRv6_medium_rec_safetensors
Purpose: Handwritten and typewritten text recognition
Languages: Swedish and Finnish
License: Apache 2.0
This is a fine-tuned model for Swedish and Finnish handwritten and typewritten text recognition, trained with computing resources generously provided by CSC – IT Center for Science on the LUMI supercomputer.
This model was developed in the ArchXAI project funded by the Central Baltic Programme.
Model Architecture
This model is a fine-tuned version of the original PP-OCRv6_medium_rec from Zhang et al (2026), for text recognition in primarily Swedish and Finnish-language historical documents.
PP-OCRv6_medium_rec is the largest recognition model in the PP-OCRv6 series. It uses LCNetV4 as the backbone and EncoderWithLightSVTR as the recognition neck, with a CTC+NRTR multi-head decoder during training and only the CTC decoder head during inference.
Intended Use
- Document digitization (e.g., archival work, historical manuscripts)
- Handwritten and typewritten notes transcription
- Data rescue (handwritten or typed numerical and letter codes)
Training data
The training data consists of human-annotated samples of mainly handwritten text lines from historical documents (16th to 20th century) in the collections of the National Archives of Finland, the Swedish National Archives and the Central Archives for Finnish Business Records (ELKA), with some printed and typed text lines included. The training set (but not the validation and test set) was expanded with empty lines containing no text, from the collections of the National Archives of Finland. A page-wise split was used for the training, validation and test sets, with the training set expanded by additional mainly numeric sources as explained below.
Training set: 1 445 496 text lines
- 1 291 832 mainly handwritten lines from NAF and RA (including 8882 empty lines from NAF)
- 153 664 mainly typewritten/printed text lines from our AIDA dataset (synthetic and from ELKA)
Validation set: 10 001 mainly handwritten text lines from NAF and RA and 4744 typewritten text lines from AIDA dataset (ELKA)
Test set (in-domain mainly HTR): 50 055 mainly handwritten text lines from NAF and RA
UoS Data Rescue: Training set (above) expanded by 6.66% with table cells sampled randomly (fresh samples per epoch) from UoS Data Rescue by Middleton and Singh (2025), a set of 768 074 table cells including some text cells in other languages such as English and French.
Randomly generated numerical table cells: Training set (above) expanded by 3.33% with digit strings randomly generated (per epoch) with our generator by concatenating individual digits from NIST SD19 (Grother and Hanaoka, 2016) and DIDA (Kusetogullari et. al., 2021).
Evaluation
The following metrics were calculated on the test set (in-domain evaluation) and two typewriting-focused test sets using the evaluate library with default settings:
HTR Test set (in-domain)
(This is the test set from the training data section)
CER (character error rate): 0.0612
WER (word error rate): 0.2527
OCR Test set 1 (NAF internal "best" typewritten only, out-of-domain)
CER (character error rate): 0.0200
WER (word error rate): 0.0797
OCR Test set 2 (AIDA "best" typewritten only, in-domain)
CER (character error rate): 0.0172
WER (word error rate): 0.0743
Used Hyperparameters
Train batch size per device (first_bs): 32
Number of devices: 128
Learning rate: 5e-5
Scheduler: cosine
Optimizer: Adam
Number of epochs: 50
Input image size: 96 x 1536
How to Use the Model
You can use the model for inference with the ppocrv6_onnx package and our PP-OCRv6 evaluation script. Note that the space character must be added to the provided dictionary, as in our evaluation script. A new inference pipeline is still in progress.
Comparison against pretrained PP-OCRv6
Stock PP-OCRv6 (PaddlePaddle's pretrained model) had the following results:
HTR Test set CER: 0.5455 WER: 0.9765
OCR Test set 1 CER: 0.0301 WER: 0.1585
OCR Test set 2 CER: 0.0303 WER: 0.1640
Limitations and Biases
The model was trained primarily on handwritten text that uses basic Latin characters and Swedish and Finnish special characters. It has not been trained on non-Latin alphabets, such as Chinese characters or other writing systems like Arabic or Hebrew. The model may not generalize well to any other language than Swedish and Finnish.
Future Work
Potential improvements for this model include:
- Expanding training data: incorporating more ground truth data
- Optimizing for specific domains: fine-tuning the model on domain-specific handwriting
- Out-of-domain generalization: studying how pre-training and fine-tuning could be optimized to maximize out-of-domain generalization of the fine-tuned model
Citation
If you use this model in your work, please cite it as:
@misc{multicentury-ppocrv6-medium,
author = {Kansallisarkisto (National Archives of Finland)},
title = {NAF Multicentury PP-OCRv6 model},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Kansallisarkisto/multicentury-ppocrv6-medium/}},
}
References
P. Grother and K. Hanaoka, NIST special database 19 handprinted forms and characters 2nd Edition, National Institute of Standards and Technology, Tech. Rep., 2016. DOI: http://doi.org/10.18434/T4H01C
Huseyin Kusetogullari, Amir Yavariabdi, Johan Hall, Niklas Lavesson, “DIGITNET: A Deep Handwritten Digit Detection and Recognition Method Using a New Historical Handwritten Digit Dataset”, Big Data Research, 2020, DOI: http://doi.org/10.1016/j.bdr.2020.100182.
Huseyin Kusetogullari, Amir Yavariabdi, Johan Hall, Niklas Lavesson, DIDA: The largest historical handwritten digit dataset with 250k digits, June 2021. Accessed on: June 13, 2021. Available: https://github.com/didadataset/DIDA/.
Middleton, S. E., & Singh, L. G. (2025). UoS Data Rescue [Dataset]. In International Journal on Document Analysis and Recognition (Version 1.0). Zenodo. https://doi.org/10.5281/zenodo.15730546
Yubo Zhang, Xueqing Wang, Manhui Lin, Yue Zhang, Penglongyi Deng, Ting Sun, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Changda Zhou, Hongen Liu, Suyin Liang, Cheng Cui, Yi Liu, Dianhai Yu, Yanjun Ma, “PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks”, arXiv preprint arXiv:2606.13108 [cs.CV], 2026. Available: https://arxiv.org/abs/2606.13108
Model Card Authors
Author: Kansallisarkisto
Contact Information: john.makela@kansallisarkisto.fi, ilkka.jokipii@kansallisarkisto.fi
- Downloads last month
- -
Model tree for Kansallisarkisto/multicentury-ppocrv6-medium
Base model
PaddlePaddle/PP-OCRv6_medium_rec_safetensors