TibetanCodexAITeam's picture
Add PechaBridge line-segmentation model card
6d1698d verified
|
Raw
History Blame Contribute Delete
3.96 kB
metadata
language:
  - bo
license: agpl-3.0
library_name: ultralytics
pipeline_tag: image-segmentation
datasets:
  - openpecha/OCR-Tibetan_line_segmentation_coordinate_annotation
tags:
  - tibetan
  - document-ai
  - line-segmentation
  - instance-segmentation
  - yolo11
  - ultralytics

PechaBridgeLineSegmentation

YOLO11n instance-segmentation model for detecting individual Tibetan text lines on traditional pecha page scans. The model is the line-layout stage used by PechaBridge before OCR.

Important: The supported PechaBridge inference pipeline applies the gray preprocessing mode (minimum RGB channel, no binarization) before YOLO inference. Raw Ultralytics inference can produce different results.

Recommended usage — PechaBridge CLI

git clone https://github.com/CodexAITeam/PechaBridge.git
cd PechaBridge
pip install -r requirements.txt

# Downloads this model and the PechaBridge OCR model.
python cli.py download-models

python cli.py batch-ocr \
    --input-dir /path/to/pecha/pages \
    --ocr-model models/ocr/PechaBridgeOCR \
    --line-model models/line_segmentation/PechaBridgeLineSegmentation.pt \
    --layout-engine yolo_line \
    --line-preprocess gray \
    --ocr-engine donut

Python usage

from pathlib import Path

from huggingface_hub import snapshot_download
from PIL import Image
from ultralytics import YOLO

from pechabridge.ocr.line_segmentation import (
    apply_line_segmentation_preprocess,
)

model_dir = Path(snapshot_download(
    "TibetanCodexAITeam/PechaBridgeLineSegmentation"
))
weights = next(model_dir.glob("*.pt"))
model = YOLO(str(weights))

image = Image.open("page.jpg").convert("RGB")
prepared = apply_line_segmentation_preprocess(image, pipeline="gray")
results = model.predict(
    source=prepared,
    imgsz=1280,
    conf=0.25,
    verbose=False,
)

for result in results:
    print(result.boxes.xyxy)  # line bounding boxes
    print(result.masks.xy)    # line polygons

Model details

  • Architecture: Ultralytics YOLO11n segmentation
  • Task: single-class instance segmentation
  • Class: line
  • Weights: yolo_line_seg.pt (about 6.0 MB)
  • Recommended input size: 1280
  • Recommended confidence threshold: 0.25
  • Recommended preprocessing: PechaBridge gray
  • Ultralytics version used for the exported checkpoint: 8.4.14

Training data

The training corpus was derived from openpecha/OCR-Tibetan_line_segmentation_coordinate_annotation, converted to YOLO polygons, padded, filtered, and split locally:

Split Images Label files
Train 3,487 3,487
Validation 519 519
Test 503 503

Validation results

Metrics stored in the released checkpoint:

Output Precision Recall mAP50 mAP50–95
Bounding boxes 0.97736 0.96768 0.99138 0.61298
Segmentation masks 0.90993 0.90055 0.89697 0.41885

These numbers describe the local validation split and should not be treated as an independent cross-collection benchmark.

Intended use and limitations

  • Intended for locating horizontal Tibetan text lines in traditional pecha scans.
  • Not an OCR model; detected line crops must be passed to a text recognizer.
  • Performance may degrade on modern book layouts, handwriting, vertical text, heavy page curvature, severe blur, unusual ornaments, or unseen collections.
  • Closely spaced or touching lines may be merged; damaged lines may be fragmented.
  • Predictions should be reviewed before scholarly or archival publication.

License

AGPL-3.0. The checkpoint was trained with Ultralytics YOLO, whose open-source software and trained model weights are distributed under AGPL-3.0 by default. Commercial or closed-source use may require an Ultralytics Enterprise License.