Bina 0.2 - RizehPizeh

Bina 0.2 - RizehPizeh is a compact Persian line-recognition model fine-tuned from PP-OCRv6 medium recognition. It recognizes both handwritten and printed Persian text using a 161-symbol character dictionary and a CTC decoder.

This repository preserves the earlier experimental checkpoints under checkpoints/. The verified release is at the repository root.

Verified release

Held-out split Rows Exact lines Exact-line accuracy Mean normalized edit similarity
Combined 4,096 3,474 84.81% 98.81%
Handwriting 2,048 1,758 85.84% 98.83%
Printed 2,048 1,716 83.79% 98.79%

Exact-line accuracy is strict: one incorrect character, digit, space, or punctuation mark makes the entire line incorrect. The independent evaluator accepted all 4,096 transforms without replacement.

Package contents

  • inference/: exported Paddle inference graph and parameters.
  • bina_text_recognition.py: logical-order Persian inference wrapper.
  • Bina-0.2-RizehPizeh.pdparams: best training checkpoint.
  • config.yml: portable architecture and preprocessing configuration.
  • persian_arabic_bina02_dict.txt: 161-symbol character dictionary.
  • MODEL_PROVENANCE.json: revisions, counts, metrics, and SHA-256 hashes.
  • eval-aligned-row-analysis-v2.json: independent row-level evaluation.

Inference

Install the appropriate PaddlePaddle build for your hardware, then:

pip install "paddleocr>=3.3.0,<4"
git clone https://huggingface.co/Reza2kn/Bina-0.2-RizehPizeh
cd Bina-0.2-RizehPizeh
python bina_text_recognition.py path/to/line.jpg --device cpu

Or use it from Python:

from bina_text_recognition import BinaTextRecognition

model = BinaTextRecognition("inference", device="cpu")
for result in model.predict("line.jpg"):
    print(result["text"], result["score"])

The exported CTC model emits Persian in visual order. Use the included wrapper to recover logical reading order while preserving Latin and numeric runs.

Input contract

The recognizer expects a horizontal, content-tight line crop:

  • Rotate tall detector crops 90 degrees clockwise.
  • Remove excessive printed-image background while retaining a small margin.
  • Resize to 3 x 48 x 768 with padding: false.
  • Preserve ZWNJ and Persian digits.

For full-page OCR, pair this recognizer with PaddlePaddle/PP-OCRv6_medium_det_safetensors, rectify and normalize each detected crop, recognize each line, then reconstruct the page in reading order. This repository contains the recognizer, not the detector.

Training

The verified release used 169,000 balanced training lines:

  • 84,500 teacher-aligned handwriting crops.
  • 84,500 printed line images.
  • 4,096 page-isolated held-out evaluation lines.

Training stopped automatically after validation plateaued. The best checkpoint was selected around step 47,520. Dataset revisions and preprocessing details are pinned in MODEL_PROVENANCE.json.

Limitations

  • Persian-focused; other-language performance was not preserved or evaluated.
  • Full-page use requires a separate detector and reading-order reconstruction.
  • The handwriting split is teacher-aligned and may contain residual label noise.
  • Very small isolated symbols and unusually degraded crops remain difficult.

License

Apache License 2.0. See LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Reza2kn/Bina-0.2-RizehPizeh

Finetuned
(1)
this model

Datasets used to train Reza2kn/Bina-0.2-RizehPizeh