中文 | English

# DocLayout Page Number Detection A YOLO-based dataset and model for detecting page numbers in document images. ## Introduction This dataset is built from the [DocLayNet-v1.2](https://huggingface.co/datasets/ds4sd/DocLayNet-v1.2) parquet files: - 14,184 images selected from `misc/selected_pngs.txt` were annotated. See that file for the corresponding parquet list. Some images share similar page number layouts and were not all annotated; others have no page numbers. - 4,899 images annotated in total: `box_single/` (4,816 images) and `box_multi/` (83 images). Note: decide whether to use `box_multi/`, as it may mislead the model. - 334 images without page numbers included as negative examples under `no_boxes/`. - 1,026 images from Chinese-format books selected from [HuggingFace](https://huggingface.co/datasets/wqzh/Gaokao-Admission-Chart-Segment) or [ModelScope](https://www.modelscope.cn/datasets/wqzh117/Gaokao-Admission-Chart-Segment) are under `pipe-adp/`. ## Labels ``` ├── DocLayout-PageNumDet/ ├── labels/ │ ├── box_single/ │ │ ├── image_1.txt │ │ ├── image_2.txt │ │ │ ├── no_boxes/ │ ├── image_a.txt (empty txt file) │ ├── image_b.txt ... ``` ## Models ``` DocLayout-PageNumDet/models ├── ch_PP-OCRv3_rec_infer.onnx (PaddleOCR recognition model) ├── ppocr_keys_v1.txt (OCR character set) ├── yolo11-medium-page-best.pt (YOLO11 medium, .pt format) ├── yolo11-medium-page-best-quant-ir9.onnx (ONNX quantized) ├── yolo11-nano-page-best.pt ├── yolo11-nano-page-best-quant-ir9.onnx ├── yolo11-small-page-best.pt └── yolo11-small-page-best-quant-ir9.onnx ``` For better detection performance, consider model ensemble. ## Usage Process a single image: detect page number boxes, run OCR, and return: ``` [[x1, y1, x2, y2, conf, text], ...] ``` - List length = number of detected boxes. - Each tuple = box coordinates, confidence, and recognized text (page number). - Cropped page number images are saved in `./tmp/`. ```shell # python src/det_and_ocr_pt.py # python src/det_and_ocr_onnx.py # python src/det_and_ocr_onnx_wo_opencv.py # without OpenCV dependency # single image CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_pt.py examples/00-80T-80_433.png CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_onnx.py examples/00-80T-80_433.png # directory CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_pt.py examples/ CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_onnx.py examples/ ``` ## ONNX Export Export `.pt` models to ONNX with quantization: ```shell python src/export_onnx.py ``` ### ONNX File Size Comparison ``` 37M models/yolo11-small-page-best-ir9.onnx 37M models/yolo11-small-page-best.onnx 37M models/yolo11-small-page-best.onnx.data 19M models/yolo11-small-page-best.pt 9.3M models/yolo11-small-page-best-quant-ir9.onnx ```