中文 | English
# DocLayout Page Number Detection
A YOLO-based dataset and model for detecting page numbers in document images.
## Introduction
This dataset is built from the [DocLayNet-v1.2](https://huggingface.co/datasets/ds4sd/DocLayNet-v1.2) parquet files:
- 14,184 images selected from `misc/selected_pngs.txt` were annotated. See that file for the corresponding parquet list. Some images share similar page number layouts and were not all annotated; others have no page numbers.
- 4,899 images annotated in total: `box_single/` (4,816 images) and `box_multi/` (83 images). Note: decide whether to use `box_multi/`, as it may mislead the model.
- 334 images without page numbers included as negative examples under `no_boxes/`.
- 1,026 images from Chinese-format books selected from [HuggingFace](https://huggingface.co/datasets/wqzh/Gaokao-Admission-Chart-Segment) or [ModelScope](https://www.modelscope.cn/datasets/wqzh117/Gaokao-Admission-Chart-Segment) are under `pipe-adp/`.
## Labels
```
├── DocLayout-PageNumDet/
├── labels/
│ ├── box_single/
│ │ ├── image_1.txt
│ │ ├── image_2.txt
│ │
│ ├── no_boxes/
│ ├── image_a.txt (empty txt file)
│ ├── image_b.txt
...
```
## Models
```
DocLayout-PageNumDet/models
├── ch_PP-OCRv3_rec_infer.onnx (PaddleOCR recognition model)
├── ppocr_keys_v1.txt (OCR character set)
├── yolo11-medium-page-best.pt (YOLO11 medium, .pt format)
├── yolo11-medium-page-best-quant-ir9.onnx (ONNX quantized)
├── yolo11-nano-page-best.pt
├── yolo11-nano-page-best-quant-ir9.onnx
├── yolo11-small-page-best.pt
└── yolo11-small-page-best-quant-ir9.onnx
```
For better detection performance, consider model ensemble.
## Usage
Process a single image: detect page number boxes, run OCR, and return:
```
[[x1, y1, x2, y2, conf, text], ...]
```
- List length = number of detected boxes.
- Each tuple = box coordinates, confidence, and recognized text (page number).
- Cropped page number images are saved in `./tmp/`.
```shell
# python src/det_and_ocr_pt.py
# python src/det_and_ocr_onnx.py
# python src/det_and_ocr_onnx_wo_opencv.py # without OpenCV dependency
# single image
CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_pt.py examples/00-80T-80_433.png
CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_onnx.py examples/00-80T-80_433.png
# directory
CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_pt.py examples/
CUDA_VISIBLE_DEVICES=1 python src/det_and_ocr_onnx.py examples/
```
## ONNX Export
Export `.pt` models to ONNX with quantization:
```shell
python src/export_onnx.py
```
### ONNX File Size Comparison
```
37M models/yolo11-small-page-best-ir9.onnx
37M models/yolo11-small-page-best.onnx
37M models/yolo11-small-page-best.onnx.data
19M models/yolo11-small-page-best.pt
9.3M models/yolo11-small-page-best-quant-ir9.onnx
```