PiotrSty/trocr-pl-mixed-v3 (experimental)

Longest fine-tune of PiotrSty/trocr-pl-base on synthetic Polish print + real EHRI typewritten Polish lines (CC-BY 4.0, ehri-pl-lines). Same recipe as trocr-pl-mixed-v2 (run 4) but 30 epochs instead of 15, because run 4's val CER was still improving at epoch 15.

Training

  • Base: PiotrSty/trocr-pl-base
  • Method: QLoRA on decoder attention (q/k/v/out_proj), rank 16, alpha 32
  • Train: 2000 synthetic lines + 349 real EHRI lines (3 docs)
  • Val: 38 EHRI lines (held-out doc ZIH3010905)
  • Epochs: 30, batch 8, lr 1e-4, T4 x2
  • Best checkpoint: /kaggle/working/trocr-pl-run5/checkpoint-4116 (val CER 0.2976, WER 0.6684)
  • Document-level split, no line leakage.

Evaluation on frozen held-out sets

Model EHRI test (81, typewriter) real-lines-v1 (75, print)
trocr-pl-base (run2) CER 47.30% / WER 90.82% CER 11.11% / WER 35.84%
trocr-pl-mixed-v1 (run3, 5 ep) CER 33.95% / WER 85.69% CER 7.09% / WER 29.44%
trocr-pl-mixed-v2 (run4, 15 ep) CER 30.42% / WER 78.85% CER 5.47% / WER 23.20%
trocr-pl-mixed-v3 (run5, 30 ep) (fill from eval cell) (fill from eval cell)

Limitations

  • Line recognizer only; page segmentation on faded typewriter is unreliable.
  • Only 349 real typewritten training lines.
  • Do NOT use as drop-in replacement without your own eval.

Provenance

See run.json, selection.json, best_metrics.json in this repo. Source: https://github.com/PiotrStyla/OCR_engine (commit c39baf4) EHRI dataset: https://huggingface.co/datasets/PiotrSty/ehri-pl-lines


Line-band experiment: completed development diagnostic

Date: 2026-09-24. Model weights were not changed or retrained. Result archive: body-dev-diagnostic-20260924T104545380613Z.zip. Previous rectangular run: body-dev-diagnostic-20260924T102032899332Z.zip. Machine-readable metrics and source hashes: summary JSON.

Verified result

Both pinned models completed 126 predictions: 63 original crops and 63 candidate/fallback crops. No inference errors or empty outputs. Five payload checksums verified; metrics recomputed from predictions and draft references. The mixed-v3 original-arm prediction records are identical to the previous run.

Model Original CER Rectangle CER Line-band CER
PiotrSty/trocr-pl-mixed-v3 31.6441% 29.4391% 29.4004%
microsoft/trocr-base-printed 85.2998% 84.4874% 84.4101%

All columns use the same 63 IDs and draft reference strings. Primary CER preserves case and historical spelling (including acute accents and long s); only NFC and whitespace normalization is applied.

Mixed-v3 WER remains 77.5656% in both candidate variants, versus 79.4749% on original crops. The new variant gains only one character edit over the rectangle variant (0.0387 percentage points CER), and 58 over original crops. Against rectangles, nine lines improve, ten regress and 44 tie in edit count. This is not evidence of a meaningful general improvement.

The new variant changes 33 crops, not the previous 31: its changed-only CER must not be directly compared to the old changed-only score. Within this run, the changed-33 subset improves from 45.0947% to 40.1033% CER; 20 improve, nine regress and four tie. All 30 fallback predictions are identical across arms. The 56-line subset excluding needs-review references worsens slightly versus rectangles: 29.0295% to 29.1161% CER.

For Microsoft, the lowercase-only diagnostic changes from 37.2147% (rectangle) to 37.2921% (band), a slight regression despite the primary CER improvement. Do not infer a clean quality win from its case-sensitive score alone.

Specific observations

  • Wiesc_FT__436884__r001__line002: terminal colon appears in the new OCR, but total character edits remain 18, versus 15 on the original crop.
  • Choragiew_FT__436804__r002__line000: 22 to 17 edits after masking, versus 19 on the original crop.
  • Slawna_wiktoria_FT__437103__r003__line003: still one error, losing the historical acute accent. The reference is not modernized to hide that error.

Scope and decision

These are draft user-corrected references from six pages, not approved gold labels. Repeated review was skipped; some references remain uncertain or have private-use glyphs. Geometry changes were developed after examining this same sample. This is not independent full-page evaluation or proof of SOTA, and it does not establish upstream training-set separation.

Do not promote band masking to the default or select per-line variants using these scores. Stop iterative tuning on these 63 lines. The next quality check should use separately selected development-validation pages, with a frozen candidate and unchanged evaluation policy; keep the final frozen test untouched.

Reproduction

  • Colab notebook
  • Input method and limitations
  • Model revisions: mixed-v3 85d0c91c26f8e088849096dded7c9ba10b4cd9c9; Microsoft 93450be3f1ed40a930690d951ef3932687cc1892.
  • Transformers 4.57.6, huggingface_hub 0.36.0, jiwer 4.0.0, Pillow 11.3.0, Torch 2.11.0+cu128; greedy FP32 decoding on Tesla T4.

Published artifacts contain aggregate metrics and provenance, not new weights or raw OCR predictions. The result ZIPs remain local.


Independent geometry holdout: rectangle retained

Date: 2026-09-24. Model weights were not changed or retrained. A geometry candidate frozen before OCR was evaluated on 12 regions from 12 previously unused development collections. All regions remained in the denominator.

Geometry CER WER Lowercase CER diagnostic
Rectangle 37.9231% 84.4444% 37.1199%
Line-band 38.0092% 83.0769% 37.3207%

Line-band produced three more character edits overall: five regions improved and seven regressed. Paired bootstrap resampling over regions gives a 95% percentile interval of [-1.2055, +1.3597] percentage points for line-band minus rectangle CER. The interval crosses zero, so rectangle remains the default and this holdout is closed for geometry tuning.

This is a geometry development holdout, not a model benchmark. The 12-region sample contains 45 private-use reference characters in 11 regions, and the upstream references have not been manually adjudicated. Primary CER uses NFC plus whitespace normalization only. Historical á, long ſ and other source glyphs are preserved and remain meaningful model errors.

The mixed-v3 rectangle error profile contains 1,322 character edits. Of these, 943 are general glyph/recognition errors, 84 are diacritic-related, 83 involve long ſ/s, 71 involve whitespace, 70 punctuation, 45 private-use reference characters and 26 case only. These diagnostic categories do not alter scoring.

Next work should focus on scan-grounded reference review and a recognizer whose training alphabet covers historical Polish print, rather than additional line-band tuning.

No source scans or raw OCR predictions from this holdout are published.

Downloads last month
127
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PiotrSty/trocr-pl-mixed-v3

Finetuned
(4)
this model