jstAnotherCapi commited on
Commit
b75331f
·
verified ·
1 Parent(s): ef6b62a

Upload OCR_testing.ipynb with huggingface_hub

Browse files
Files changed (1) hide show
  1. OCR_testing.ipynb +2419 -0
OCR_testing.ipynb ADDED
@@ -0,0 +1,2419 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "cells": [
3
+ {
4
+ "cell_type": "markdown",
5
+ "id": "511312c4-f049-415b-91c0-bc54ebd6e5c7",
6
+ "metadata": {
7
+ "jp-MarkdownHeadingCollapsed": true
8
+ },
9
+ "source": [
10
+ "# Test 1 : qari-ocr"
11
+ ]
12
+ },
13
+ {
14
+ "cell_type": "code",
15
+ "execution_count": 7,
16
+ "id": "d71883d5-65b0-4b89-bf53-8e9c9bef6b2e",
17
+ "metadata": {},
18
+ "outputs": [],
19
+ "source": [
20
+ "import torch\n",
21
+ "from transformers import Qwen2VLForConditionalGeneration, AutoProcessor\n",
22
+ "from peft import PeftModel\n",
23
+ "from PIL import Image\n",
24
+ "import os\n",
25
+ "from qwen_vl_utils import process_vision_info"
26
+ ]
27
+ },
28
+ {
29
+ "cell_type": "code",
30
+ "execution_count": 8,
31
+ "id": "afaddd27-0a56-4ee0-8c0a-e0c7e3354c89",
32
+ "metadata": {},
33
+ "outputs": [],
34
+ "source": [
35
+ "\n",
36
+ "\n",
37
+ "# Paths to your local directories\n",
38
+ "BASE_MODEL_PATH = \"./Qwen2-VL-2B-Instruct\" # Update this\n",
39
+ "ADAPTER_PATH = \"./qari-ocr\" # Your current adapter directory"
40
+ ]
41
+ },
42
+ {
43
+ "cell_type": "code",
44
+ "execution_count": 3,
45
+ "id": "30a29f8d-889b-4552-bbfc-d98093d00338",
46
+ "metadata": {},
47
+ "outputs": [],
48
+ "source": [
49
+ "# Disable transformers from trying to connect to the internet\n",
50
+ "os.environ['HF_HUB_OFFLINE'] = '1'\n",
51
+ "os.environ['TRANSFORMERS_OFFLINE'] = '1'"
52
+ ]
53
+ },
54
+ {
55
+ "cell_type": "code",
56
+ "execution_count": 4,
57
+ "id": "a6196e9a-a173-48a4-a9f9-486afa425cbd",
58
+ "metadata": {},
59
+ "outputs": [
60
+ {
61
+ "name": "stdout",
62
+ "output_type": "stream",
63
+ "text": [
64
+ "Loading model and processor...\n"
65
+ ]
66
+ }
67
+ ],
68
+ "source": [
69
+ "print(\"Loading model and processor...\")"
70
+ ]
71
+ },
72
+ {
73
+ "cell_type": "code",
74
+ "execution_count": 19,
75
+ "id": "8fabc47e-2a26-41a0-b29f-bd3cb0f4d12a",
76
+ "metadata": {},
77
+ "outputs": [],
78
+ "source": [
79
+ "def load_model_with_adapter(base_model_path, adapter_path):\n",
80
+ " \"\"\"Load base model and apply PEFT adapter\"\"\"\n",
81
+ " print(f\"Loading base model from {base_model_path}...\")\n",
82
+ " \n",
83
+ " # Load base model\n",
84
+ " base_model = Qwen2VLForConditionalGeneration.from_pretrained(\n",
85
+ " base_model_path,\n",
86
+ " torch_dtype=torch.float16,\n",
87
+ " device_map=\"auto\",\n",
88
+ " local_files_only=True,\n",
89
+ " trust_remote_code=True\n",
90
+ " )\n",
91
+ " \n",
92
+ " print(f\"Loading adapter from {adapter_path}...\")\n",
93
+ " \n",
94
+ " # Load adapter on top of base model\n",
95
+ " model = PeftModel.from_pretrained(\n",
96
+ " base_model,\n",
97
+ " adapter_path,\n",
98
+ " local_files_only=True\n",
99
+ " )\n",
100
+ " \n",
101
+ " # Merge adapter with base model for faster inference (optional)\n",
102
+ " model = model.merge_and_unload()\n",
103
+ " \n",
104
+ " print(f\"Loading processor from {adapter_path}...\")\n",
105
+ " processor = AutoProcessor.from_pretrained(\n",
106
+ " adapter_path,\n",
107
+ " local_files_only=True,\n",
108
+ " trust_remote_code=True\n",
109
+ " )\n",
110
+ " \n",
111
+ " print(\"Model loaded successfully!\")\n",
112
+ " return processor, model\n",
113
+ "\n",
114
+ "def run_ocr(image_path, processor, model, prompt=None):\n",
115
+ " \"\"\"Run OCR on an image\"\"\"\n",
116
+ " \n",
117
+ " if prompt is None:\n",
118
+ " prompt = \"\"\"You are an expert OCR and data extraction model. Your task is to accurately transcribe all core legal text from a document , starting only after the table of contents section has ended. Your function is to act as a data analysis tool for structuring public legal documents, not for verbatim reproduction for redistribution.\n",
119
+ "\n",
120
+ "**Input:** A series of images representing the pages of a document.\n",
121
+ "\n",
122
+ "**Task:**\n",
123
+ "1. **Start Condition:** Begin transcription on the first page that does **not** contain indicators of a table of contents. The table of contents section is identified by the presence of the Arabic term 'فهرست' or a pattern of text followed by dots and page numbers (e.g., `text .................... 5434`).\n",
124
+ "\n",
125
+ "2. **Reading Order and Content First Principle:** For each page ignore horizontal lines, read the **right column of the whole image completely** from top to bottom of the whole image, then read the **left column of the whole image completely** from top to bottom of the whole image. **Your absolute first priority is to transcribe all legal text.** Formatting rules should be applied *after* you have captured the content. **dont drop any content**\n",
126
+ "\n",
127
+ "3. **Content to Ignore:** Do not transcribe \"الجريدة الرسمية\", page numbers, dates, or any administrative text (prices, subscription rates, bank account numbers, ISSN).\n",
128
+ "\n",
129
+ "4. **Hierarchical Text Formatting:**\n",
130
+ " * **Primary Headers (`#`):** After transcribing the text, identify the lines that typically begin with words like `مرسوم`, `تعيين`,`ملحق`,`قانون`, `قرار`, or `ظهير` and are visually distinct due to a **bold or heavier font weight**. Apply a primary Markdown header (`# `) **only** to these specific, bolded title lines. \n",
131
+ " * **Secondary Headers (`##`):** Identify structural sub-divisions (`الباب`, `الفصل`, `المادة`). For lines that begin **directly** with these words, use a secondary Markdown header (`## `). **Do not** apply this header if these keywords are preceded by other characters on the same line (e.g., do not format `«المادة...` as a header).These headers can be used independently and do not need to follow a primary header. Insert a blank line before and after any valid header\n",
132
+ " \n",
133
+ "\n",
134
+ "5. **Table Formatting:**\n",
135
+ " * Represent tables using **Markdown table syntax**.\n",
136
+ "\n",
137
+ "6. **Paragraph and Line Break Rules (in order of priority):**\n",
138
+ " * **Always insert a new line** after the phrases `قرر ما يلي :` or `رسم ما يلي :`.\n",
139
+ " * **Always start on a new line** for any paragraph that begins with `المادة`.\n",
140
+ " * **For improved readability in preambles:** In the introductory section of a law (the text before `قرر ما يلي :`), insert a line break after each semi-colon (`;`) that separates different legal citations or clauses.\n",
141
+ " * Insert a line break when a block of Arabic text is followed by a block of Latin text (or vice versa). This does not apply to numbers within a sentence.\n",
142
+ " * For all other text, combine lines into clean, continuous paragraphs, breaking only at sentence-ending punctuation.\n",
143
+ "\n",
144
+ "7. **Stop Condition:** Extract all content until the end of the document.\n",
145
+ "\n",
146
+ "8. **Conditional Extraction:** Return \" \" (an empty string) for any page identified as part of the table of contents.\n",
147
+ "\n",
148
+ "**Output Format:**\n",
149
+ "Provide the extracted text as a single, continuous block of Markdown. The output should be a clean, well-structured, and hierarchically organized transcription of the legal content. Do not hallucinate.\"\"\"\n",
150
+ " \n",
151
+ " # Load image\n",
152
+ " image = Image.open(image_path).convert(\"RGB\")\n",
153
+ " \n",
154
+ " # Prepare messages in the expected format\n",
155
+ " messages = [\n",
156
+ " {\n",
157
+ " \"role\": \"user\",\n",
158
+ " \"content\": [\n",
159
+ " {\"type\": \"image\", \"image\": image},\n",
160
+ " {\"type\": \"text\", \"text\": prompt},\n",
161
+ " ],\n",
162
+ " }\n",
163
+ " ]\n",
164
+ " \n",
165
+ " # Apply chat template\n",
166
+ " text = processor.apply_chat_template(\n",
167
+ " messages, \n",
168
+ " tokenize=False, \n",
169
+ " add_generation_prompt=True\n",
170
+ " )\n",
171
+ " \n",
172
+ " # Process vision info\n",
173
+ " image_inputs, video_inputs = process_vision_info(messages)\n",
174
+ " \n",
175
+ " # Prepare inputs\n",
176
+ " inputs = processor(\n",
177
+ " text=[text],\n",
178
+ " images=image_inputs,\n",
179
+ " videos=video_inputs,\n",
180
+ " padding=True,\n",
181
+ " return_tensors=\"pt\",\n",
182
+ " )\n",
183
+ " inputs = inputs.to(model.device)\n",
184
+ " \n",
185
+ " # Generate output\n",
186
+ " print(\"Running OCR...\")\n",
187
+ " with torch.no_grad():\n",
188
+ " generated_ids = model.generate(\n",
189
+ " **inputs, \n",
190
+ " max_new_tokens=2000\n",
191
+ " )\n",
192
+ " \n",
193
+ " # Trim the input tokens from output\n",
194
+ " generated_ids_trimmed = [\n",
195
+ " out_ids[len(in_ids):] \n",
196
+ " for in_ids, out_ids in zip(inputs.input_ids, generated_ids)\n",
197
+ " ]\n",
198
+ " \n",
199
+ " # Decode output\n",
200
+ " output_text = processor.batch_decode(\n",
201
+ " generated_ids_trimmed, \n",
202
+ " skip_special_tokens=True, \n",
203
+ " clean_up_tokenization_spaces=False\n",
204
+ " )[0]\n",
205
+ " \n",
206
+ " return output_text\n",
207
+ "\n"
208
+ ]
209
+ },
210
+ {
211
+ "cell_type": "code",
212
+ "execution_count": 15,
213
+ "id": "09396de0-936c-487a-9511-277cc4fb993e",
214
+ "metadata": {},
215
+ "outputs": [],
216
+ "source": [
217
+ "if not os.path.exists(BASE_MODEL_PATH):\n",
218
+ " raise FileNotFoundError(\n",
219
+ " f\"Model directory not found at {BASE_MODEL_PATH}. \"\n",
220
+ " \"Please update MODEL_PATH variable.\")\n",
221
+ " \n"
222
+ ]
223
+ },
224
+ {
225
+ "cell_type": "code",
226
+ "execution_count": 10,
227
+ "id": "d1a0f957-4328-4981-bbeb-463730a9fa02",
228
+ "metadata": {},
229
+ "outputs": [
230
+ {
231
+ "name": "stderr",
232
+ "output_type": "stream",
233
+ "text": [
234
+ "`torch_dtype` is deprecated! Use `dtype` instead!\n"
235
+ ]
236
+ },
237
+ {
238
+ "name": "stdout",
239
+ "output_type": "stream",
240
+ "text": [
241
+ "Loading base model from ./Qwen2-VL-2B-Instruct...\n"
242
+ ]
243
+ },
244
+ {
245
+ "name": "stderr",
246
+ "output_type": "stream",
247
+ "text": [
248
+ "Loading checkpoint shards: 100%|██████████| 2/2 [00:00<00:00, 2.47it/s]\n"
249
+ ]
250
+ },
251
+ {
252
+ "name": "stdout",
253
+ "output_type": "stream",
254
+ "text": [
255
+ "Loading adapter from ./qari-ocr...\n"
256
+ ]
257
+ },
258
+ {
259
+ "name": "stderr",
260
+ "output_type": "stream",
261
+ "text": [
262
+ "The image processor of type `Qwen2VLImageProcessor` is now loaded as a fast processor by default, even if the model checkpoint was saved with a slow processor. This is a breaking change and may produce slightly different outputs. To continue using the slow processor, instantiate this class with `use_fast=False`. Note that this behavior will be extended to all models in a future release.\n"
263
+ ]
264
+ },
265
+ {
266
+ "name": "stdout",
267
+ "output_type": "stream",
268
+ "text": [
269
+ "Loading processor from ./qari-ocr...\n",
270
+ "Model loaded successfully!\n"
271
+ ]
272
+ }
273
+ ],
274
+ "source": [
275
+ "processor, model = load_model_with_adapter(BASE_MODEL_PATH, ADAPTER_PATH)"
276
+ ]
277
+ },
278
+ {
279
+ "cell_type": "code",
280
+ "execution_count": 21,
281
+ "id": "ae8b5de6-5d2b-4e04-b0a3-a8e41fb84336",
282
+ "metadata": {
283
+ "scrolled": true
284
+ },
285
+ "outputs": [
286
+ {
287
+ "name": "stdout",
288
+ "output_type": "stream",
289
+ "text": [
290
+ "Running OCR...\n",
291
+ "\n",
292
+ "==================================================\n",
293
+ "OCR Result:\n",
294
+ "==================================================\n",
295
+ "قرر ما يلي :\n",
296
+ "المادة الأولى\n",
297
+ ": Génie civil\n",
298
+ "تقبل لمعادلة دبلوم مهندس دولة، الشهادة التالية في Mines-Télécom Lille Douai, de l’Institut Mines-Télécom,\n",
299
+ "délivré en date du 10 janvier 2024 - France.\n",
300
+ "المادة الثانية\n",
301
+ "ينشر هذا القرار بالجريدة الرسمية.\n",
302
+ "وزير التعليم العالي والبحث العلمي والابتكار رقم 2061.24\n",
303
+ "قرار لوزير التعليم العالي والبحث العلمي والابتكار رقم 1446 (25 يوليوز 2024) بتحديد بعض المعادلات بين الشهادات.\n",
304
+ "وزير التعليم العالي والبحث العلمي والابتكار، بناء على المرسوم رقم 2.01.333 الصادر في 28 من ربيع الأول 1422\n",
305
+ "(21 يونيوز 2001) المتعلق بتحديد الشروط والمسطرة الخاصة بمنح معادلة شهادات التعليم العالي ;\n",
306
+ "و على المرسوم رقم 2.21.838 الصادر في 14 من ربيع الأول 1443\n",
307
+ "(21 أكتوبر 2021) المتعلق باختصاصات وزير التعليم العالي والبحث العلمي والابتكار ;\n",
308
+ "وبعد استشارة اللجنة القطاعية للعلوم والتقنيات والهندسة والهندسة المعمارية المنعقدة بتاريخ 4 يوليوز 2024،\n",
309
+ "قرر ما يلي :\n",
310
+ "المادة الأولى\n",
311
+ ": Informatique\n",
312
+ "تقبل لمعادلة دبلوم مهندس دولة، الشهادة التالية في Mines-Télécom Nancy de l’Université de Lorraine, délivré en date du 26 octobre 2023 - France.\n"
313
+ ]
314
+ }
315
+ ],
316
+ "source": [
317
+ "# Example: Run OCR on a single image\n",
318
+ "image_path = \"journal_25_right.png\" # Update with your image path\n",
319
+ " \n",
320
+ "if os.path.exists(image_path):\n",
321
+ " result = run_ocr(image_path, processor, model)\n",
322
+ " print(\"\\n\" + \"=\"*50)\n",
323
+ " print(\"OCR Result:\")\n",
324
+ " print(\"=\"*50)\n",
325
+ " print(result)\n",
326
+ " \n",
327
+ " # Example: Batch process directory\n",
328
+ " # Uncomment to use batch processing\n",
329
+ " # batch_process_images(\"./images\", processor, model, \"results.txt\")\n",
330
+ "else:\n",
331
+ " print(f\"No test image found at {image_path}\")\n",
332
+ " print(\"Please provide an image path to test.\") "
333
+ ]
334
+ },
335
+ {
336
+ "cell_type": "code",
337
+ "execution_count": 22,
338
+ "id": "74637485-5094-4b29-8359-d9bccfb38ad4",
339
+ "metadata": {
340
+ "scrolled": true
341
+ },
342
+ "outputs": [
343
+ {
344
+ "name": "stdout",
345
+ "output_type": "stream",
346
+ "text": [
347
+ "Running OCR...\n",
348
+ "\n",
349
+ "==================================================\n",
350
+ "OCR Result:\n",
351
+ "==================================================\n",
352
+ "6. وبعد معالجة ودراسة جميع الشكايات والطلبات، تبين أن ما يناهز 994 شكاية وطلب تدخل في إطار اختصاصات المجلس تمت معالجتها واتخاذ ما يلزم بخصوصها، بالإضافة إلى 115 شكاية عالجتها الآليتين الوطنيتين الخاصتين بحقوق الطفل والأشخاص في وضعية اعاقة، في حين تبين أن 1913 شكاية وطلب لا تندرج ضمن اختصاص المجلس ولجانه وآلياته تم توجيه المعنين بها إلى سلك المساطر القانونية أو الإدارية، أو تمت إحالتها على الجهات المختصة، بما في ذلك 134 شكاية تمت إحالتها على مؤسسة وسيط المملكة للاختصاص، فيما تم حفظ 296 شكاية لكونها لا تحترم الشروط القانونية والواقعية لقبولها أو مجهولة المصدر، أو سبق البت فيها.\n",
353
+ "7. وحسب التصنيف الموضوعاتي لهذه الشكايات، بلغ عدد الشكايات المتعلقة بالحقوق المدنية والسياسية ما يناهز 343 شكاية وطلب، منها 152 شكاية تتعلق بادعاءات المس بالحق في السلامة الجسدية، و29 شكاية تتعلق بحرية الجمعيات والعمل النقابي، في حين توزعت باقي الشكايات البالغ عددها 162 على الحقوق الأخرى، ومن بينها الحق في التجمع والتظاهر وحرية الرأي والتعبير والحق في المحاكمة العادلة. أما فيما يخص الحقوق الاقتصادية والاجتماعية والثقافية والبيئية، فقد بلغ عدد الشكايات بشأنها ما مجموعه 651 شكاية وطلب.\n",
354
+ "8. وحسب التصنيف انطلاقا من حقوق بعض الفئات والمجموعات، فقد تلقى المجلس ولجانه الجهوية 276 شكاية تهم حقوق المهاجرين، و280 شكاية من نساء أو فتيات ضحايا العنف، في حين بلغ عدد الشكايات والطلبات الواردة من سجناء أو ذويهم 1312 شكاية وطلب، تتوزع على طلبات العفو والتظلم من الأحكام القضائية وطلبات الترحيل أو الاحتفاظ بنفس المؤسسة السجنية وإعادة التصنيف، وتظلمات تهم التطبيق ومتابعة الدراسة والاتصال بالعالم الخارجي وادعاءات سوء المعاملة.\n",
355
+ "9. وقد عمل المجلس على دراسة هذه الشكايات جميعها دراسة دقيقة للوقوف على حقيقة الادعاءات الواردة فيها، وأحال ما يقتضي إحالتها على الجهات المعنية للتحري في موضوع الانتهاكات المحتملة، مع متابعتها وصولا الى معالجتها وإشعار المشتكين بكل الإجراءات المتخذة بخصوصها. كما تم إرشاد بعضهم وتوجيههم إلى سلوك المساطر والإجراءات القانونية المخول لهم اتباعها لطرح أو متابعة إجراءات معالجة شكاياتهم أو تظلماتهم أمام الجهات المختصة باعتبارها هي وحدها المعنية بالبت فيها وكذا تبليغهم بالقرارات المتخذة بشأنها.\n",
356
+ "10. ويسجل المجلس تفاعل القطاعات الحكومية مع الشكايات التي يحيلها عليها، إلا أن هذا التفاعل ما زال يتم بدرجات تتفاوت بين قطاع وآخر، كما أن نوعية الأجوبة تبقى أغلبها ذات طبيعة عامة وتبريرية، مما يجعلها غير مقنعة بالنسبة لموضوع الادعاء. كما يسجل المجلس، في الكثير من الحالات، عدم\n",
357
+ "\n",
358
+ "✅ OCR text saved to: journal_7_ocr.txt\n"
359
+ ]
360
+ }
361
+ ],
362
+ "source": [
363
+ "# Example: Run OCR on a single image\n",
364
+ "image_path = \"journal_7.png\" # Update with your image path\n",
365
+ "output_txt_path = \"journal_7_ocr.txt\" # where to save the result\n",
366
+ "\n",
367
+ "if os.path.exists(image_path):\n",
368
+ " result = run_ocr(image_path, processor, model)\n",
369
+ " print(\"\\n\" + \"=\"*50)\n",
370
+ " print(\"OCR Result:\")\n",
371
+ " print(\"=\"*50)\n",
372
+ " print(result)\n",
373
+ "\n",
374
+ " # ---> Save OCR result to TXT file <---\n",
375
+ " with open(output_txt_path, \"w\", encoding=\"utf-8\") as f:\n",
376
+ " f.write(result)\n",
377
+ " print(f\"\\n✅ OCR text saved to: {output_txt_path}\")\n",
378
+ "\n",
379
+ " # Example: Batch process directory\n",
380
+ " # Uncomment to use batch processing\n",
381
+ " # batch_process_images(\"./images\", processor, model, \"results.txt\")\n",
382
+ "else:\n",
383
+ " print(f\"No test image found at {image_path}\")\n",
384
+ " print(\"Please provide an image path to test.\")\n"
385
+ ]
386
+ },
387
+ {
388
+ "cell_type": "code",
389
+ "execution_count": 23,
390
+ "id": "c9de868c-2725-41c2-80ea-c6f74cfd3c34",
391
+ "metadata": {},
392
+ "outputs": [],
393
+ "source": [
394
+ "import os, glob, re, datetime\n",
395
+ "\n",
396
+ "def _natural_sort_key(path):\n",
397
+ " \"\"\"\n",
398
+ " Natural sort: splits digits so 'page_2.png' < 'page_10.png'.\n",
399
+ " \"\"\"\n",
400
+ " name = os.path.basename(path)\n",
401
+ " return [int(t) if t.isdigit() else t.lower() for t in re.findall(r'\\d+|\\D+', name)]\n",
402
+ "\n",
403
+ "def ocr_all_png_to_one(images_dir, output_txt_path, processor, model, recursive=False):\n",
404
+ " \"\"\"\n",
405
+ " Runs OCR on all PNGs in images_dir and writes the combined result into output_txt_path.\n",
406
+ " \n",
407
+ " Parameters\n",
408
+ " ----------\n",
409
+ " images_dir : str\n",
410
+ " Directory to scan for .png files.\n",
411
+ " output_txt_path : str\n",
412
+ " Destination .txt file to write all OCR results into (overwrites).\n",
413
+ " processor, model :\n",
414
+ " Your existing OCR components used by run_ocr(image_path, processor, model).\n",
415
+ " recursive : bool\n",
416
+ " If True, search subfolders as well.\n",
417
+ " \"\"\"\n",
418
+ " pattern = \"**/*.png\" if recursive else \"*.png\"\n",
419
+ " files = glob.glob(os.path.join(images_dir, pattern), recursive=recursive)\n",
420
+ " files = [f for f in files if os.path.isfile(f)]\n",
421
+ " if not files:\n",
422
+ " print(f\"⚠️ No PNG files found in: {images_dir}\")\n",
423
+ " return\n",
424
+ "\n",
425
+ " files.sort(key=_natural_sort_key)\n",
426
+ " ts = datetime.datetime.now().strftime(\"%Y-%m-%d %H:%M:%S\")\n",
427
+ "\n",
428
+ " count_ok, count_err = 0, 0\n",
429
+ " with open(output_txt_path, \"w\", encoding=\"utf-8\") as out:\n",
430
+ " out.write(f\"# Combined OCR output\\n# Source folder: {os.path.abspath(images_dir)}\\n# Generated: {ts}\\n\\n\")\n",
431
+ "\n",
432
+ " for i, img_path in enumerate(files, 1):\n",
433
+ " header = f\"=== File {i}/{len(files)}: {os.path.basename(img_path)} ===\"\n",
434
+ " print(f\"OCR: {header}\")\n",
435
+ " out.write(header + \"\\n\")\n",
436
+ "\n",
437
+ " try:\n",
438
+ " text = run_ocr(img_path, processor, model) or \"\"\n",
439
+ " out.write(text.strip() + \"\\n\\n\")\n",
440
+ " count_ok += 1\n",
441
+ " except Exception as e:\n",
442
+ " msg = f\"[ERROR processing {img_path}: {e}]\"\n",
443
+ " print(msg)\n",
444
+ " out.write(msg + \"\\n\\n\")\n",
445
+ " count_err += 1\n",
446
+ "\n",
447
+ " print(f\"✅ Done. Wrote {count_ok} files (errors: {count_err}) to: {output_txt_path}\")\n"
448
+ ]
449
+ },
450
+ {
451
+ "cell_type": "code",
452
+ "execution_count": 25,
453
+ "id": "960a647e-d8cd-4b11-8e96-cefa9fb92999",
454
+ "metadata": {
455
+ "scrolled": true
456
+ },
457
+ "outputs": [
458
+ {
459
+ "name": "stdout",
460
+ "output_type": "stream",
461
+ "text": [
462
+ "OCR: === File 1/32: journal_1_left.png ===\n",
463
+ "Running OCR...\n",
464
+ "OCR: === File 2/32: journal_1_right.png ===\n",
465
+ "Running OCR...\n",
466
+ "OCR: === File 3/32: journal_2.png ===\n",
467
+ "Running OCR...\n",
468
+ "OCR: === File 4/32: journal_3.png ===\n",
469
+ "Running OCR...\n",
470
+ "OCR: === File 5/32: journal_4.png ===\n",
471
+ "Running OCR...\n",
472
+ "OCR: === File 6/32: journal_5.png ===\n",
473
+ "Running OCR...\n",
474
+ "OCR: === File 7/32: journal_6.png ===\n",
475
+ "Running OCR...\n",
476
+ "OCR: === File 8/32: journal_7.png ===\n",
477
+ "Running OCR...\n",
478
+ "OCR: === File 9/32: journal_8.png ===\n",
479
+ "Running OCR...\n",
480
+ "OCR: === File 10/32: journal_9.png ===\n",
481
+ "Running OCR...\n",
482
+ "OCR: === File 11/32: journal_10.png ===\n",
483
+ "Running OCR...\n",
484
+ "OCR: === File 12/32: journal_12.png ===\n",
485
+ "Running OCR...\n",
486
+ "OCR: === File 13/32: journal_13.png ===\n",
487
+ "Running OCR...\n",
488
+ "OCR: === File 14/32: journal_14.png ===\n",
489
+ "Running OCR...\n",
490
+ "OCR: === File 15/32: journal_15_left.png ===\n",
491
+ "Running OCR...\n",
492
+ "OCR: === File 16/32: journal_15_right.png ===\n",
493
+ "Running OCR...\n",
494
+ "OCR: === File 17/32: journal_16_left.png ===\n",
495
+ "Running OCR...\n",
496
+ "OCR: === File 18/32: journal_16_right.png ===\n",
497
+ "Running OCR...\n",
498
+ "OCR: === File 19/32: journal_17_left.png ===\n",
499
+ "Running OCR...\n",
500
+ "OCR: === File 20/32: journal_17_right.png ===\n",
501
+ "Running OCR...\n",
502
+ "OCR: === File 21/32: journal_18_left.png ===\n",
503
+ "Running OCR...\n",
504
+ "OCR: === File 22/32: journal_18_right.png ===\n",
505
+ "Running OCR...\n",
506
+ "OCR: === File 23/32: journal_19.png ===\n",
507
+ "Running OCR...\n",
508
+ "OCR: === File 24/32: journal_20.png ===\n",
509
+ "Running OCR...\n",
510
+ "OCR: === File 25/32: journal_21.png ===\n",
511
+ "Running OCR...\n",
512
+ "OCR: === File 26/32: journal_22.png ===\n",
513
+ "Running OCR...\n",
514
+ "OCR: === File 27/32: journal_23.png ===\n",
515
+ "Running OCR...\n",
516
+ "OCR: === File 28/32: journal_24_left.png ===\n",
517
+ "Running OCR...\n",
518
+ "OCR: === File 29/32: journal_24_right.png ===\n",
519
+ "Running OCR...\n",
520
+ "OCR: === File 30/32: journal_25_left.png ===\n",
521
+ "Running OCR...\n",
522
+ "OCR: === File 31/32: journal_25_right.png ===\n",
523
+ "Running OCR...\n",
524
+ "OCR: === File 32/32: journal_27.png ===\n",
525
+ "Running OCR...\n",
526
+ "✅ Done. Wrote 32 files (errors: 0) to: ./output/combined_ocr.txt\n"
527
+ ]
528
+ }
529
+ ],
530
+ "source": [
531
+ "# Set your folder and output file\n",
532
+ "images_dir = \"./images\" # folder containing your .png files\n",
533
+ "output_txt = \"./output/output_qari-ocr.txt\" # single output file\n",
534
+ "\n",
535
+ "ocr_all_png_to_one(images_dir, output_txt, processor, model, recursive=False)\n"
536
+ ]
537
+ },
538
+ {
539
+ "cell_type": "markdown",
540
+ "id": "3999fb8d-0539-4678-9c52-18f28f51455e",
541
+ "metadata": {},
542
+ "source": [
543
+ "# Test 2 : Nanonets-OCR2-3B"
544
+ ]
545
+ },
546
+ {
547
+ "cell_type": "markdown",
548
+ "id": "883c1658-b826-4091-8a12-5f831eac6f67",
549
+ "metadata": {},
550
+ "source": [
551
+ "## with my prompt and one single image "
552
+ ]
553
+ },
554
+ {
555
+ "cell_type": "code",
556
+ "execution_count": 1,
557
+ "id": "c8c38588-69cd-4938-8c63-e4e68ad222a7",
558
+ "metadata": {},
559
+ "outputs": [
560
+ {
561
+ "name": "stderr",
562
+ "output_type": "stream",
563
+ "text": [
564
+ "/home/skiredj.abderrahman/.conda/envs/gptoss-vllm/lib/python3.10/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html\n",
565
+ " from .autonotebook import tqdm as notebook_tqdm\n"
566
+ ]
567
+ }
568
+ ],
569
+ "source": [
570
+ "from PIL import Image\n",
571
+ "from transformers import AutoTokenizer, AutoProcessor, AutoModelForImageTextToText\n"
572
+ ]
573
+ },
574
+ {
575
+ "cell_type": "code",
576
+ "execution_count": 2,
577
+ "id": "f6c07569-26c3-4f52-ab6a-485c286f0bf8",
578
+ "metadata": {},
579
+ "outputs": [
580
+ {
581
+ "ename": "NameError",
582
+ "evalue": "name 'os' is not defined",
583
+ "output_type": "error",
584
+ "traceback": [
585
+ "\u001b[0;31m---------------------------------------------------------------------------\u001b[0m",
586
+ "\u001b[0;31mNameError\u001b[0m Traceback (most recent call last)",
587
+ "Cell \u001b[0;32mIn[2], line 1\u001b[0m\n\u001b[0;32m----> 1\u001b[0m \u001b[43mos\u001b[49m\u001b[38;5;241m.\u001b[39menviron[\u001b[38;5;124m\"\u001b[39m\u001b[38;5;124mCUDA_VISIBLE_DEVICES\u001b[39m\u001b[38;5;124m\"\u001b[39m] \u001b[38;5;241m=\u001b[39m \u001b[38;5;124m'\u001b[39m\u001b[38;5;124m0\u001b[39m\u001b[38;5;124m'\u001b[39m\n",
588
+ "\u001b[0;31mNameError\u001b[0m: name 'os' is not defined"
589
+ ]
590
+ }
591
+ ],
592
+ "source": [
593
+ "os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0'"
594
+ ]
595
+ },
596
+ {
597
+ "cell_type": "code",
598
+ "execution_count": 47,
599
+ "id": "5f429dbe-2054-4786-a227-829f67933bd0",
600
+ "metadata": {
601
+ "scrolled": true
602
+ },
603
+ "outputs": [
604
+ {
605
+ "name": "stderr",
606
+ "output_type": "stream",
607
+ "text": [
608
+ "Loading checkpoint shards: 100%|██████████| 2/2 [00:01<00:00, 1.61it/s]\n"
609
+ ]
610
+ },
611
+ {
612
+ "data": {
613
+ "text/plain": [
614
+ "Qwen2_5_VLForConditionalGeneration(\n",
615
+ " (model): Qwen2_5_VLModel(\n",
616
+ " (visual): Qwen2_5_VisionTransformerPretrainedModel(\n",
617
+ " (patch_embed): Qwen2_5_VisionPatchEmbed(\n",
618
+ " (proj): Conv3d(3, 1280, kernel_size=(2, 14, 14), stride=(2, 14, 14), bias=False)\n",
619
+ " )\n",
620
+ " (rotary_pos_emb): Qwen2_5_VisionRotaryEmbedding()\n",
621
+ " (blocks): ModuleList(\n",
622
+ " (0-31): 32 x Qwen2_5_VLVisionBlock(\n",
623
+ " (norm1): Qwen2RMSNorm((1280,), eps=1e-06)\n",
624
+ " (norm2): Qwen2RMSNorm((1280,), eps=1e-06)\n",
625
+ " (attn): Qwen2_5_VLVisionAttention(\n",
626
+ " (qkv): Linear(in_features=1280, out_features=3840, bias=True)\n",
627
+ " (proj): Linear(in_features=1280, out_features=1280, bias=True)\n",
628
+ " )\n",
629
+ " (mlp): Qwen2_5_VLMLP(\n",
630
+ " (gate_proj): Linear(in_features=1280, out_features=3420, bias=True)\n",
631
+ " (up_proj): Linear(in_features=1280, out_features=3420, bias=True)\n",
632
+ " (down_proj): Linear(in_features=3420, out_features=1280, bias=True)\n",
633
+ " (act_fn): SiLUActivation()\n",
634
+ " )\n",
635
+ " )\n",
636
+ " )\n",
637
+ " (merger): Qwen2_5_VLPatchMerger(\n",
638
+ " (ln_q): Qwen2RMSNorm((1280,), eps=1e-06)\n",
639
+ " (mlp): Sequential(\n",
640
+ " (0): Linear(in_features=5120, out_features=5120, bias=True)\n",
641
+ " (1): GELU(approximate='none')\n",
642
+ " (2): Linear(in_features=5120, out_features=2048, bias=True)\n",
643
+ " )\n",
644
+ " )\n",
645
+ " )\n",
646
+ " (language_model): Qwen2_5_VLTextModel(\n",
647
+ " (embed_tokens): Embedding(151936, 2048)\n",
648
+ " (layers): ModuleList(\n",
649
+ " (0-35): 36 x Qwen2_5_VLDecoderLayer(\n",
650
+ " (self_attn): Qwen2_5_VLAttention(\n",
651
+ " (q_proj): Linear(in_features=2048, out_features=2048, bias=True)\n",
652
+ " (k_proj): Linear(in_features=2048, out_features=256, bias=True)\n",
653
+ " (v_proj): Linear(in_features=2048, out_features=256, bias=True)\n",
654
+ " (o_proj): Linear(in_features=2048, out_features=2048, bias=False)\n",
655
+ " (rotary_emb): Qwen2_5_VLRotaryEmbedding()\n",
656
+ " )\n",
657
+ " (mlp): Qwen2MLP(\n",
658
+ " (gate_proj): Linear(in_features=2048, out_features=11008, bias=False)\n",
659
+ " (up_proj): Linear(in_features=2048, out_features=11008, bias=False)\n",
660
+ " (down_proj): Linear(in_features=11008, out_features=2048, bias=False)\n",
661
+ " (act_fn): SiLUActivation()\n",
662
+ " )\n",
663
+ " (input_layernorm): Qwen2RMSNorm((2048,), eps=1e-06)\n",
664
+ " (post_attention_layernorm): Qwen2RMSNorm((2048,), eps=1e-06)\n",
665
+ " )\n",
666
+ " )\n",
667
+ " (norm): Qwen2RMSNorm((2048,), eps=1e-06)\n",
668
+ " (rotary_emb): Qwen2_5_VLRotaryEmbedding()\n",
669
+ " )\n",
670
+ " )\n",
671
+ " (lm_head): Linear(in_features=2048, out_features=151936, bias=False)\n",
672
+ ")"
673
+ ]
674
+ },
675
+ "execution_count": 47,
676
+ "metadata": {},
677
+ "output_type": "execute_result"
678
+ }
679
+ ],
680
+ "source": [
681
+ "model_path = \"Nanonets-OCR2-3B\"\n",
682
+ "\n",
683
+ "model = AutoModelForImageTextToText.from_pretrained(\n",
684
+ " model_path, \n",
685
+ " #bf 16\n",
686
+ " torch_dtype=\"auto\", \n",
687
+ " device_map=\"auto\", \n",
688
+ " \n",
689
+ ")\n",
690
+ "model.eval()"
691
+ ]
692
+ },
693
+ {
694
+ "cell_type": "code",
695
+ "execution_count": 48,
696
+ "id": "9aeaae0f-eb58-4f11-9710-e9010c8e78e5",
697
+ "metadata": {},
698
+ "outputs": [],
699
+ "source": [
700
+ "tokenizer = AutoTokenizer.from_pretrained(model_path)\n",
701
+ "processor = AutoProcessor.from_pretrained(model_path)\n"
702
+ ]
703
+ },
704
+ {
705
+ "cell_type": "code",
706
+ "execution_count": 49,
707
+ "id": "5e612e64-e5f3-4f46-b666-664098bfe403",
708
+ "metadata": {},
709
+ "outputs": [],
710
+ "source": [
711
+ "\n",
712
+ "def ocr_page_with_nanonets_s(image_path, model, processor, max_new_tokens=4096):\n",
713
+ " prompt = \"\"\"Extract the text from the above document as if you were reading it naturally. Return the tables in html format. Return the equations in LaTeX representation. If there is an image in the document and image caption is not present, add a small description of the image inside the <img></img> tag; otherwise, add the image caption inside <img></img>. Watermarks should be wrapped in brackets. Ex: <watermark>OFFICIAL COPY</watermark>. Page numbers should be wrapped in brackets. Ex: <page_number>14</page_number> or <page_number>9/22</page_number>. Prefer using ☐ and ☑ for check boxes.\"\"\"\n",
714
+ " image = Image.open(image_path)\n",
715
+ " messages = [\n",
716
+ " {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n",
717
+ " {\"role\": \"user\", \"content\": [\n",
718
+ " {\"type\": \"image\", \"image\": f\"file://{image_path}\"},\n",
719
+ " {\"type\": \"text\", \"text\": prompt},\n",
720
+ " ]},\n",
721
+ " ]\n",
722
+ " text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)\n",
723
+ " inputs = processor(text=[text], images=[image], padding=True, return_tensors=\"pt\")\n",
724
+ " inputs = inputs.to(model.device)\n",
725
+ " \n",
726
+ " output_ids = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)\n",
727
+ " generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]\n",
728
+ " \n",
729
+ " output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)\n",
730
+ " return output_text[0]"
731
+ ]
732
+ },
733
+ {
734
+ "cell_type": "code",
735
+ "execution_count": 50,
736
+ "id": "009b9b8d-6ee2-47e9-9234-b137f58d2f06",
737
+ "metadata": {},
738
+ "outputs": [
739
+ {
740
+ "name": "stdout",
741
+ "output_type": "stream",
742
+ "text": [
743
+ "6402\n",
744
+ "الجريدة\n",
745
+ "\n",
746
+ "قرر ما يلي:\n",
747
+ "المادة الأولى\n",
748
+ "تقبل معادلة دبلوم مهندس دولة، الشهادة التالية في : Génie civil\n",
749
+ "- Titre d'ingénieur diplôme de l'Ecole nationale supérieure Mines-Télécom Lille Douai, de l'Institut Mines-Télécom, délivré en date du 10 janvier 2024 - France.\n",
750
+ "\n",
751
+ "المادة الثانية\n",
752
+ "ينشر هذا القرار بالجريدة الرسمية.\n",
753
+ "وحرر بالرباط في 19 من محرم 1446 (25 يوليو 2024).\n",
754
+ "الإمضاء: عبد اللطيف ميراوي.\n",
755
+ "\n",
756
+ "قرار لوزير التعليم العالي والبحث العلمي والابتكار رقم 2061.24 صادر في 19 من محرم 1446 (25 يوليو 2024) بتحديد بعض المعادلات بين الشهادات.\n",
757
+ "\n",
758
+ "وزير التعليم العالي والبحث العلمي والابتكار، بناء على المرسوم رقم 333.01 الصادر في 28 من ربيع الأول 1422 (21 يونيو 2001) المتعلق بتحديد الشروط والمسطرة الخاصة بمنح معادلة شهادات التعليم العالي؛ وعلى المرسوم رقم 89.04.2 الصادر في 18 من ربيع الآخر 1425 (7 يونيو 2004) المتعلق باختصاص المؤسسات الجامعية وأسلاك الدراسات العليا وكذا الشهادات الوطنية المطابقة، كما وقع تغييره وتتميمه؛ وعلى المرسوم رقم 838.21 الصادر في 14 من ربيع الأول 1443 (21 أكتوبر 2021) المتعلق باختصاصات وزير التعليم العالي والبحث العلمي والابتكار؛ وبعد استشارة اللجنة القطاعية للعلوم والتقنيات والهندسة والهندسة المعمارية المنعقدة بتاريخ 4 يوليو 2024، قرر ما يلي:\n",
759
+ "المادة الأولى\n",
760
+ ": Informatique\n",
761
+ "تقبل معادلة دبلوم مهندس دولة، الشهادة التالية في :\n",
762
+ "- Titre d'ingénieur de télécom Nancy de l'Université de Lorraine, délivré en date du 26 octobre 2023 - France.\n"
763
+ ]
764
+ }
765
+ ],
766
+ "source": [
767
+ "image_path = \"journal_25_right.png\" # Change this\n",
768
+ "result = ocr_page_with_nanonets_s(image_path, model, processor, max_new_tokens=15000)\n",
769
+ "print(result)"
770
+ ]
771
+ },
772
+ {
773
+ "cell_type": "code",
774
+ "execution_count": 3,
775
+ "id": "09d59e5a-6758-4fcf-a962-124a5ac50acd",
776
+ "metadata": {},
777
+ "outputs": [
778
+ {
779
+ "name": "stderr",
780
+ "output_type": "stream",
781
+ "text": [
782
+ "`torch_dtype` is deprecated! Use `dtype` instead!\n"
783
+ ]
784
+ },
785
+ {
786
+ "name": "stdout",
787
+ "output_type": "stream",
788
+ "text": [
789
+ "Loading model...\n"
790
+ ]
791
+ },
792
+ {
793
+ "name": "stderr",
794
+ "output_type": "stream",
795
+ "text": [
796
+ "Loading checkpoint shards: 100%|██████████| 2/2 [00:01<00:00, 1.57it/s]\n"
797
+ ]
798
+ },
799
+ {
800
+ "name": "stdout",
801
+ "output_type": "stream",
802
+ "text": [
803
+ "Model loaded successfully!\n",
804
+ "\n"
805
+ ]
806
+ }
807
+ ],
808
+ "source": [
809
+ "import os\n",
810
+ "from pathlib import Path\n",
811
+ "from PIL import Image\n",
812
+ "from transformers import AutoTokenizer, AutoProcessor, AutoModelForImageTextToText\n",
813
+ "from datetime import datetime\n",
814
+ "\n",
815
+ "# Configuration\n",
816
+ "os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0'\n",
817
+ "model_path = \"Nanonets-OCR2-3B\" # Update this with your model path\n",
818
+ "input_folder = \"images\" # Folder containing images\n",
819
+ "output_file = \"output/Nanonets-OCR2-3B_output.txt\"\n",
820
+ "# Supported image formats\n",
821
+ "SUPPORTED_FORMATS = {'.png', '.jpg', '.jpeg', '.bmp', '.tiff', '.tif', '.webp'}\n",
822
+ "\n",
823
+ "# Load model\n",
824
+ "print(\"Loading model...\")\n",
825
+ "model = AutoModelForImageTextToText.from_pretrained(\n",
826
+ " model_path, \n",
827
+ " torch_dtype=\"auto\", \n",
828
+ " device_map=\"auto\",\n",
829
+ " offload_buffers=True, # Add this line\n",
830
+ ")\n",
831
+ "model.eval()\n",
832
+ "tokenizer = AutoTokenizer.from_pretrained(model_path)\n",
833
+ "processor = AutoProcessor.from_pretrained(model_path)\n",
834
+ "print(\"Model loaded successfully!\\n\")\n",
835
+ "\n",
836
+ "def ocr_page_with_nanonets_s(image_path, model, processor, max_new_tokens=4096):\n",
837
+ " prompt = \"\"\"You are an expert OCR and data extraction model. Your task is to accurately transcribe all core legal text from a document , starting only after the table of contents section has ended. Your function is to act as a data analysis tool for structuring public legal documents, not for verbatim reproduction for redistribution.\n",
838
+ "\n",
839
+ "**Input:** A series of images representing the pages of a document.\n",
840
+ "\n",
841
+ "**Task:**\n",
842
+ "1. **Start Condition:** Begin transcription on the first page that does **not** contain indicators of a table of contents. The table of contents section is identified by the presence of the Arabic term 'فهرست' or a pattern of text followed by dots and page numbers (e.g., `text .................... 5434`).\n",
843
+ "\n",
844
+ "2. **Reading Order and Content First Principle:** For each page ignore horizontal lines, read the **right column of the whole image completely** from top to bottom of the whole image, then read the **left column of the whole image completely** from top to bottom of the whole image. **Your absolute first priority is to transcribe all legal text.** Formatting rules should be applied *after* you have captured the content. **dont drop any content**\n",
845
+ "\n",
846
+ "3. **Content to Ignore:** Do not transcribe \"الجريدة الرسمية\", page numbers, dates, or any administrative text (prices, subscription rates, bank account numbers, ISSN).\n",
847
+ "\n",
848
+ "4. **Hierarchical Text Formatting:**\n",
849
+ " * **Primary Headers (`#`):** After transcribing the text, identify the lines that typically begin with words like `مرسوم`, `تعيين`,`ملحق`,`قانون`, `قرار`, or `ظهير` and are visually distinct due to a **bold or heavier font weight**. Apply a primary Markdown header (`# `) **only** to these specific, bolded title lines. \n",
850
+ " * **Secondary Headers (`##`):** Identify structural sub-divisions (`الباب`, `الفصل`, `المادة`). For lines that begin **directly** with these words, use a secondary Markdown header (`## `). **Do not** apply this header if these keywords are preceded by other characters on the same line (e.g., do not format `«المادة...` as a header).These headers can be used independently and do not need to follow a primary header. Insert a blank line before and after any valid header\n",
851
+ " \n",
852
+ "\n",
853
+ "5. **Table Formatting:**\n",
854
+ " * Represent tables using **Markdown table syntax**.\n",
855
+ "\n",
856
+ "6. **Paragraph and Line Break Rules (in order of priority):**\n",
857
+ " * **Always insert a new line** after the phrases `قرر ما يلي :` or `رسم ما يلي :`.\n",
858
+ " * **Always start on a new line** for any paragraph that begins with `المادة`.\n",
859
+ " * **For improved readability in preambles:** In the introductory section of a law (the text before `قرر ما يلي :`), insert a line break after each semi-colon (`;`) that separates different legal citations or clauses.\n",
860
+ " * Insert a line break when a block of Arabic text is followed by a block of Latin text (or vice versa). This does not apply to numbers within a sentence.\n",
861
+ " * For all other text, combine lines into clean, continuous paragraphs, breaking only at sentence-ending punctuation.\n",
862
+ "\n",
863
+ "7. **Stop Condition:** Extract all content until the end of the document.\n",
864
+ "\n",
865
+ "8. **Conditional Extraction:** Return \" \" (an empty string) for any page identified as part of the table of contents.\n",
866
+ "\n",
867
+ "**Output Format:**\n",
868
+ "Provide the extracted text as a single, continuous block of Markdown. The output should be a clean, well-structured, and hierarchically organized transcription of the legal content. Do not hallucinate.\"\"\"\n",
869
+ " \n",
870
+ " image = Image.open(image_path)\n",
871
+ " messages = [\n",
872
+ " {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n",
873
+ " {\"role\": \"user\", \"content\": [\n",
874
+ " {\"type\": \"image\", \"image\": f\"file://{image_path}\"},\n",
875
+ " {\"type\": \"text\", \"text\": prompt},\n",
876
+ " ]},\n",
877
+ " ]\n",
878
+ " text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)\n",
879
+ " inputs = processor(text=[text], images=[image], padding=True, return_tensors=\"pt\")\n",
880
+ " inputs = inputs.to(model.device)\n",
881
+ " \n",
882
+ " output_ids = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)\n",
883
+ " generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]\n",
884
+ " \n",
885
+ " output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)\n",
886
+ " return output_text[0]\n",
887
+ "\n",
888
+ "def process_folder(input_folder, output_file, model, processor, max_new_tokens=15000):\n",
889
+ " # Get all image files from folder\n",
890
+ " image_files = []\n",
891
+ " for ext in SUPPORTED_FORMATS:\n",
892
+ " image_files.extend(Path(input_folder).glob(f\"*{ext}\"))\n",
893
+ " image_files.extend(Path(input_folder).glob(f\"*{ext.upper()}\"))\n",
894
+ " \n",
895
+ " # Sort files by name\n",
896
+ " image_files = sorted(image_files)\n",
897
+ " \n",
898
+ " if not image_files:\n",
899
+ " print(f\"No images found in {input_folder}\")\n",
900
+ " return\n",
901
+ " \n",
902
+ " print(f\"Found {len(image_files)} images to process\\n\")\n",
903
+ " \n",
904
+ " # Process each image and write to output file\n",
905
+ " with open(output_file, 'w', encoding='utf-8') as f:\n",
906
+ " f.write(f\"OCR Results - Generated on {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\\n\")\n",
907
+ " f.write(\"=\" * 80 + \"\\n\\n\")\n",
908
+ " \n",
909
+ " for idx, image_path in enumerate(image_files, 1):\n",
910
+ " print(f\"Processing [{idx}/{len(image_files)}]: {image_path.name}\")\n",
911
+ " \n",
912
+ " try:\n",
913
+ " result = ocr_page_with_nanonets_s(str(image_path), model, processor, max_new_tokens)\n",
914
+ " \n",
915
+ " # Write to file\n",
916
+ " f.write(f\"{'=' * 80}\\n\")\n",
917
+ " f.write(f\"File: {image_path.name}\\n\")\n",
918
+ " f.write(f\"{'=' * 80}\\n\\n\")\n",
919
+ " f.write(result)\n",
920
+ " f.write(f\"\\n\\n{'=' * 80}\\n\\n\")\n",
921
+ " \n",
922
+ " print(f\"✓ Completed: {image_path.name}\\n\")\n",
923
+ " \n",
924
+ " except Exception as e:\n",
925
+ " error_msg = f\"✗ Error processing {image_path.name}: {str(e)}\"\n",
926
+ " print(error_msg)\n",
927
+ " f.write(f\"\\n{error_msg}\\n\\n\")\n",
928
+ " \n",
929
+ " print(f\"\\n{'=' * 80}\")\n",
930
+ " print(f\"Processing complete! Results saved to: {output_file}\")\n",
931
+ " print(f\"{'=' * 80}\")\n"
932
+ ]
933
+ },
934
+ {
935
+ "cell_type": "code",
936
+ "execution_count": 4,
937
+ "id": "7344fd3a-4461-400a-966c-619536ab2183",
938
+ "metadata": {
939
+ "scrolled": true
940
+ },
941
+ "outputs": [
942
+ {
943
+ "name": "stderr",
944
+ "output_type": "stream",
945
+ "text": [
946
+ "The following generation flags are not valid and may be ignored: ['temperature']. Set `TRANSFORMERS_VERBOSITY=info` for more details.\n"
947
+ ]
948
+ },
949
+ {
950
+ "name": "stdout",
951
+ "output_type": "stream",
952
+ "text": [
953
+ "Found 32 images to process\n",
954
+ "\n",
955
+ "Processing [1/32]: journal_10.png\n",
956
+ "✓ Completed: journal_10.png\n",
957
+ "\n",
958
+ "Processing [2/32]: journal_12.png\n",
959
+ "✓ Completed: journal_12.png\n",
960
+ "\n",
961
+ "Processing [3/32]: journal_13.png\n",
962
+ "✓ Completed: journal_13.png\n",
963
+ "\n",
964
+ "Processing [4/32]: journal_14.png\n",
965
+ "✓ Completed: journal_14.png\n",
966
+ "\n",
967
+ "Processing [5/32]: journal_15_left.png\n",
968
+ "✓ Completed: journal_15_left.png\n",
969
+ "\n",
970
+ "Processing [6/32]: journal_15_right.png\n",
971
+ "✓ Completed: journal_15_right.png\n",
972
+ "\n",
973
+ "Processing [7/32]: journal_16_left.png\n",
974
+ "✓ Completed: journal_16_left.png\n",
975
+ "\n",
976
+ "Processing [8/32]: journal_16_right.png\n",
977
+ "✓ Completed: journal_16_right.png\n",
978
+ "\n",
979
+ "Processing [9/32]: journal_17_left.png\n",
980
+ "✓ Completed: journal_17_left.png\n",
981
+ "\n",
982
+ "Processing [10/32]: journal_17_right.png\n",
983
+ "✓ Completed: journal_17_right.png\n",
984
+ "\n",
985
+ "Processing [11/32]: journal_18_left.png\n",
986
+ "✓ Completed: journal_18_left.png\n",
987
+ "\n",
988
+ "Processing [12/32]: journal_18_right.png\n",
989
+ "✓ Completed: journal_18_right.png\n",
990
+ "\n",
991
+ "Processing [13/32]: journal_19.png\n",
992
+ "✓ Completed: journal_19.png\n",
993
+ "\n",
994
+ "Processing [14/32]: journal_1_left.png\n",
995
+ "✓ Completed: journal_1_left.png\n",
996
+ "\n",
997
+ "Processing [15/32]: journal_1_right.png\n",
998
+ "✓ Completed: journal_1_right.png\n",
999
+ "\n",
1000
+ "Processing [16/32]: journal_2.png\n",
1001
+ "✓ Completed: journal_2.png\n",
1002
+ "\n",
1003
+ "Processing [17/32]: journal_20.png\n",
1004
+ "✓ Completed: journal_20.png\n",
1005
+ "\n",
1006
+ "Processing [18/32]: journal_21.png\n",
1007
+ "✓ Completed: journal_21.png\n",
1008
+ "\n",
1009
+ "Processing [19/32]: journal_22.png\n",
1010
+ "✓ Completed: journal_22.png\n",
1011
+ "\n",
1012
+ "Processing [20/32]: journal_23.png\n",
1013
+ "✓ Completed: journal_23.png\n",
1014
+ "\n",
1015
+ "Processing [21/32]: journal_24_left.png\n",
1016
+ "✓ Completed: journal_24_left.png\n",
1017
+ "\n",
1018
+ "Processing [22/32]: journal_24_right.png\n",
1019
+ "✓ Completed: journal_24_right.png\n",
1020
+ "\n",
1021
+ "Processing [23/32]: journal_25_left.png\n",
1022
+ "✓ Completed: journal_25_left.png\n",
1023
+ "\n",
1024
+ "Processing [24/32]: journal_25_right.png\n",
1025
+ "✓ Completed: journal_25_right.png\n",
1026
+ "\n",
1027
+ "Processing [25/32]: journal_27.png\n",
1028
+ "✓ Completed: journal_27.png\n",
1029
+ "\n",
1030
+ "Processing [26/32]: journal_3.png\n",
1031
+ "✓ Completed: journal_3.png\n",
1032
+ "\n",
1033
+ "Processing [27/32]: journal_4.png\n",
1034
+ "✓ Completed: journal_4.png\n",
1035
+ "\n",
1036
+ "Processing [28/32]: journal_5.png\n",
1037
+ "✓ Completed: journal_5.png\n",
1038
+ "\n",
1039
+ "Processing [29/32]: journal_6.png\n",
1040
+ "✓ Completed: journal_6.png\n",
1041
+ "\n",
1042
+ "Processing [30/32]: journal_7.png\n",
1043
+ "✓ Completed: journal_7.png\n",
1044
+ "\n",
1045
+ "Processing [31/32]: journal_8.png\n",
1046
+ "✓ Completed: journal_8.png\n",
1047
+ "\n",
1048
+ "Processing [32/32]: journal_9.png\n",
1049
+ "✓ Completed: journal_9.png\n",
1050
+ "\n",
1051
+ "\n",
1052
+ "================================================================================\n",
1053
+ "Processing complete! Results saved to: output/Nanonets-OCR2-3B_output.txt\n",
1054
+ "================================================================================\n"
1055
+ ]
1056
+ }
1057
+ ],
1058
+ "source": [
1059
+ "process_folder(input_folder, output_file, model, processor, max_new_tokens=15000)"
1060
+ ]
1061
+ },
1062
+ {
1063
+ "cell_type": "markdown",
1064
+ "id": "30c088c9-b6a1-44d1-bcbc-d3d43ee39fa5",
1065
+ "metadata": {},
1066
+ "source": [
1067
+ "## multiple images with skirdj "
1068
+ ]
1069
+ },
1070
+ {
1071
+ "cell_type": "code",
1072
+ "execution_count": 5,
1073
+ "id": "068df1b3-c98e-4e57-9d22-a23ae57f118f",
1074
+ "metadata": {},
1075
+ "outputs": [
1076
+ {
1077
+ "name": "stdout",
1078
+ "output_type": "stream",
1079
+ "text": [
1080
+ "Loading model...\n"
1081
+ ]
1082
+ },
1083
+ {
1084
+ "name": "stderr",
1085
+ "output_type": "stream",
1086
+ "text": [
1087
+ "Loading checkpoint shards: 100%|██████████| 2/2 [00:01<00:00, 1.72it/s]\n"
1088
+ ]
1089
+ },
1090
+ {
1091
+ "name": "stdout",
1092
+ "output_type": "stream",
1093
+ "text": [
1094
+ "Model loaded successfully!\n",
1095
+ "\n"
1096
+ ]
1097
+ }
1098
+ ],
1099
+ "source": [
1100
+ "import os\n",
1101
+ "from pathlib import Path\n",
1102
+ "from PIL import Image\n",
1103
+ "from transformers import AutoTokenizer, AutoProcessor, AutoModelForImageTextToText\n",
1104
+ "from datetime import datetime\n",
1105
+ "\n",
1106
+ "# Configuration\n",
1107
+ "os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0'\n",
1108
+ "model_path = \"Nanonets-OCR2-3B\" # Update this with your model path\n",
1109
+ "input_folder = \"images\" # Folder containing images\n",
1110
+ "output_file = \"output/Nanonets-OCR2-3B_output_skirdje_prompt.txt\"\n",
1111
+ "# Supported image formats\n",
1112
+ "SUPPORTED_FORMATS = {'.png', '.jpg', '.jpeg', '.bmp', '.tiff', '.tif', '.webp'}\n",
1113
+ "\n",
1114
+ "# Load model\n",
1115
+ "print(\"Loading model...\")\n",
1116
+ "model = AutoModelForImageTextToText.from_pretrained(\n",
1117
+ " model_path, \n",
1118
+ " torch_dtype=\"auto\", \n",
1119
+ " device_map=\"auto\",\n",
1120
+ " offload_buffers=True, # Add this line\n",
1121
+ ")\n",
1122
+ "model.eval()\n",
1123
+ "tokenizer = AutoTokenizer.from_pretrained(model_path)\n",
1124
+ "processor = AutoProcessor.from_pretrained(model_path)\n",
1125
+ "print(\"Model loaded successfully!\\n\")\n",
1126
+ "\n",
1127
+ "def ocr_page_with_nanonets_s(image_path, model, processor, max_new_tokens=4096):\n",
1128
+ " prompt = \"\"\"\n",
1129
+ "Extract all visible text from the document image, following these rules:\n",
1130
+ "\n",
1131
+ "0. *Exhaustivity*\n",
1132
+ "\n",
1133
+ " * OCR *all text faithfully*. Do not omit or alter any word or detail.\n",
1134
+ "\n",
1135
+ "1. *Line Breaks*\n",
1136
+ "\n",
1137
+ " * Preserve line breaks *only* when:\n",
1138
+ "\n",
1139
+ " * The sentence ends with a period, question mark, or exclamation mark followed by a newline.\n",
1140
+ " * A heading/title is separated from body text.\n",
1141
+ " * ALWAYS insert a line break after subtitles starting with these keywords: المادة / الفصل / الباب.\n",
1142
+ " Example:\n",
1143
+ "\n",
1144
+ " \n",
1145
+ " المادة الثانية\n",
1146
+ " الفصل 2\n",
1147
+ " الباب الثالث\n",
1148
+ " \n",
1149
+ " * In all other cases, merge text into continuous paragraphs for readability.\n",
1150
+ "\n",
1151
+ "2. *Tables*\n",
1152
+ "\n",
1153
+ " * If tables are present, reformat them into clean, valid *Markdown table syntax*.\n",
1154
+ "\n",
1155
+ "3. *Ignore Content*\n",
1156
+ "\n",
1157
+ " * Do *not* OCR:\n",
1158
+ "\n",
1159
+ " * Headers at the very top of the page (e.g., journal name like \"الجريدة الرسمية\", issue numbers, dates).\n",
1160
+ " * Footers or page numbers.\n",
1161
+ " * Any number located at the extreme bottom of the page.\n",
1162
+ "\n",
1163
+ "4. *Subtitle Marking*\n",
1164
+ "\n",
1165
+ " * Subtitles are *only* those beginning with: المادة / الفصل.\n",
1166
+ " * Prepend ## (or deeper levels if nested) before subtitles, respecting hierarchy:\n",
1167
+ "\n",
1168
+ " * ##الفصل\n",
1169
+ "\n",
1170
+ " * ###المادة\n",
1171
+ "\n",
1172
+ " * Always add a line break *after* each subtitle. So whenever you write like \"المادة الثانية\" add a linebreak after it.\n",
1173
+ " * Never use a single # heading.\n",
1174
+ "\n",
1175
+ "5. *Two-Column Pages*\n",
1176
+ "\n",
1177
+ " * If the page is laid out in *two vertical columns* (with a clear dividing line between right and left), always OCR the *right column first, then the **left column*.\n",
1178
+ " * Do *not* confuse tables with columns — this rule applies *only* to page layouts, not tables.\n",
1179
+ "\n",
1180
+ "6. *Keep all dots*\n",
1181
+ " * If you see many dots in the text you are OCRizing (like \"....\"), please keep them all. They have meanings, so do not remove them.\n",
1182
+ "\n",
1183
+ "\"\"\"\n",
1184
+ " \n",
1185
+ " image = Image.open(image_path)\n",
1186
+ " messages = [\n",
1187
+ " {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n",
1188
+ " {\"role\": \"user\", \"content\": [\n",
1189
+ " {\"type\": \"image\", \"image\": f\"file://{image_path}\"},\n",
1190
+ " {\"type\": \"text\", \"text\": prompt},\n",
1191
+ " ]},\n",
1192
+ " ]\n",
1193
+ " text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)\n",
1194
+ " inputs = processor(text=[text], images=[image], padding=True, return_tensors=\"pt\")\n",
1195
+ " inputs = inputs.to(model.device)\n",
1196
+ " \n",
1197
+ " output_ids = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)\n",
1198
+ " generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]\n",
1199
+ " \n",
1200
+ " output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)\n",
1201
+ " return output_text[0]\n",
1202
+ "\n",
1203
+ "def process_folder(input_folder, output_file, model, processor, max_new_tokens=15000):\n",
1204
+ " # Get all image files from folder\n",
1205
+ " image_files = []\n",
1206
+ " for ext in SUPPORTED_FORMATS:\n",
1207
+ " image_files.extend(Path(input_folder).glob(f\"*{ext}\"))\n",
1208
+ " image_files.extend(Path(input_folder).glob(f\"*{ext.upper()}\"))\n",
1209
+ " \n",
1210
+ " # Sort files by name\n",
1211
+ " image_files = sorted(image_files)\n",
1212
+ " \n",
1213
+ " if not image_files:\n",
1214
+ " print(f\"No images found in {input_folder}\")\n",
1215
+ " return\n",
1216
+ " \n",
1217
+ " print(f\"Found {len(image_files)} images to process\\n\")\n",
1218
+ " \n",
1219
+ " # Process each image and write to output file\n",
1220
+ " with open(output_file, 'w', encoding='utf-8') as f:\n",
1221
+ " f.write(f\"OCR Results - Generated on {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\\n\")\n",
1222
+ " f.write(\"=\" * 80 + \"\\n\\n\")\n",
1223
+ " \n",
1224
+ " for idx, image_path in enumerate(image_files, 1):\n",
1225
+ " print(f\"Processing [{idx}/{len(image_files)}]: {image_path.name}\")\n",
1226
+ " \n",
1227
+ " try:\n",
1228
+ " result = ocr_page_with_nanonets_s(str(image_path), model, processor, max_new_tokens)\n",
1229
+ " \n",
1230
+ " # Write to file\n",
1231
+ " f.write(f\"{'=' * 80}\\n\")\n",
1232
+ " f.write(f\"File: {image_path.name}\\n\")\n",
1233
+ " f.write(f\"{'=' * 80}\\n\\n\")\n",
1234
+ " f.write(result)\n",
1235
+ " f.write(f\"\\n\\n{'=' * 80}\\n\\n\")\n",
1236
+ " \n",
1237
+ " print(f\"✓ Completed: {image_path.name}\\n\")\n",
1238
+ " \n",
1239
+ " except Exception as e:\n",
1240
+ " error_msg = f\"✗ Error processing {image_path.name}: {str(e)}\"\n",
1241
+ " print(error_msg)\n",
1242
+ " f.write(f\"\\n{error_msg}\\n\\n\")\n",
1243
+ " \n",
1244
+ " print(f\"\\n{'=' * 80}\")\n",
1245
+ " print(f\"Processing complete! Results saved to: {output_file}\")\n",
1246
+ " print(f\"{'=' * 80}\")\n"
1247
+ ]
1248
+ },
1249
+ {
1250
+ "cell_type": "code",
1251
+ "execution_count": 6,
1252
+ "id": "7e5f497a-87c2-47d7-bfb4-3511cb4d5e53",
1253
+ "metadata": {
1254
+ "scrolled": true
1255
+ },
1256
+ "outputs": [
1257
+ {
1258
+ "name": "stdout",
1259
+ "output_type": "stream",
1260
+ "text": [
1261
+ "Found 32 images to process\n",
1262
+ "\n",
1263
+ "Processing [1/32]: journal_10.png\n",
1264
+ "✓ Completed: journal_10.png\n",
1265
+ "\n",
1266
+ "Processing [2/32]: journal_12.png\n",
1267
+ "✓ Completed: journal_12.png\n",
1268
+ "\n",
1269
+ "Processing [3/32]: journal_13.png\n",
1270
+ "✓ Completed: journal_13.png\n",
1271
+ "\n",
1272
+ "Processing [4/32]: journal_14.png\n",
1273
+ "✓ Completed: journal_14.png\n",
1274
+ "\n",
1275
+ "Processing [5/32]: journal_15_left.png\n",
1276
+ "✓ Completed: journal_15_left.png\n",
1277
+ "\n",
1278
+ "Processing [6/32]: journal_15_right.png\n",
1279
+ "✓ Completed: journal_15_right.png\n",
1280
+ "\n",
1281
+ "Processing [7/32]: journal_16_left.png\n",
1282
+ "✓ Completed: journal_16_left.png\n",
1283
+ "\n",
1284
+ "Processing [8/32]: journal_16_right.png\n",
1285
+ "✓ Completed: journal_16_right.png\n",
1286
+ "\n",
1287
+ "Processing [9/32]: journal_17_left.png\n",
1288
+ "✓ Completed: journal_17_left.png\n",
1289
+ "\n",
1290
+ "Processing [10/32]: journal_17_right.png\n",
1291
+ "✓ Completed: journal_17_right.png\n",
1292
+ "\n",
1293
+ "Processing [11/32]: journal_18_left.png\n",
1294
+ "✓ Completed: journal_18_left.png\n",
1295
+ "\n",
1296
+ "Processing [12/32]: journal_18_right.png\n",
1297
+ "✓ Completed: journal_18_right.png\n",
1298
+ "\n",
1299
+ "Processing [13/32]: journal_19.png\n",
1300
+ "✓ Completed: journal_19.png\n",
1301
+ "\n",
1302
+ "Processing [14/32]: journal_1_left.png\n",
1303
+ "✓ Completed: journal_1_left.png\n",
1304
+ "\n",
1305
+ "Processing [15/32]: journal_1_right.png\n",
1306
+ "✓ Completed: journal_1_right.png\n",
1307
+ "\n",
1308
+ "Processing [16/32]: journal_2.png\n",
1309
+ "✓ Completed: journal_2.png\n",
1310
+ "\n",
1311
+ "Processing [17/32]: journal_20.png\n",
1312
+ "✓ Completed: journal_20.png\n",
1313
+ "\n",
1314
+ "Processing [18/32]: journal_21.png\n",
1315
+ "✓ Completed: journal_21.png\n",
1316
+ "\n",
1317
+ "Processing [19/32]: journal_22.png\n",
1318
+ "✓ Completed: journal_22.png\n",
1319
+ "\n",
1320
+ "Processing [20/32]: journal_23.png\n",
1321
+ "✓ Completed: journal_23.png\n",
1322
+ "\n",
1323
+ "Processing [21/32]: journal_24_left.png\n",
1324
+ "✓ Completed: journal_24_left.png\n",
1325
+ "\n",
1326
+ "Processing [22/32]: journal_24_right.png\n",
1327
+ "✓ Completed: journal_24_right.png\n",
1328
+ "\n",
1329
+ "Processing [23/32]: journal_25_left.png\n",
1330
+ "✓ Completed: journal_25_left.png\n",
1331
+ "\n",
1332
+ "Processing [24/32]: journal_25_right.png\n",
1333
+ "✓ Completed: journal_25_right.png\n",
1334
+ "\n",
1335
+ "Processing [25/32]: journal_27.png\n",
1336
+ "✓ Completed: journal_27.png\n",
1337
+ "\n",
1338
+ "Processing [26/32]: journal_3.png\n",
1339
+ "✓ Completed: journal_3.png\n",
1340
+ "\n",
1341
+ "Processing [27/32]: journal_4.png\n",
1342
+ "✓ Completed: journal_4.png\n",
1343
+ "\n",
1344
+ "Processing [28/32]: journal_5.png\n",
1345
+ "✓ Completed: journal_5.png\n",
1346
+ "\n",
1347
+ "Processing [29/32]: journal_6.png\n",
1348
+ "✓ Completed: journal_6.png\n",
1349
+ "\n",
1350
+ "Processing [30/32]: journal_7.png\n",
1351
+ "✓ Completed: journal_7.png\n",
1352
+ "\n",
1353
+ "Processing [31/32]: journal_8.png\n",
1354
+ "✓ Completed: journal_8.png\n",
1355
+ "\n",
1356
+ "Processing [32/32]: journal_9.png\n",
1357
+ "✓ Completed: journal_9.png\n",
1358
+ "\n",
1359
+ "\n",
1360
+ "================================================================================\n",
1361
+ "Processing complete! Results saved to: output/Nanonets-OCR2-3B_output_skirdje_prompt.txt\n",
1362
+ "================================================================================\n"
1363
+ ]
1364
+ }
1365
+ ],
1366
+ "source": [
1367
+ "process_folder(input_folder, output_file, model, processor, max_new_tokens=15000)"
1368
+ ]
1369
+ },
1370
+ {
1371
+ "cell_type": "markdown",
1372
+ "id": "d5a7ce8f-d35b-4a13-8410-8ece9715fab8",
1373
+ "metadata": {},
1374
+ "source": [
1375
+ "## testing tables only "
1376
+ ]
1377
+ },
1378
+ {
1379
+ "cell_type": "code",
1380
+ "execution_count": 1,
1381
+ "id": "dac1d342-a84b-428a-ae27-47888537c51b",
1382
+ "metadata": {},
1383
+ "outputs": [
1384
+ {
1385
+ "name": "stderr",
1386
+ "output_type": "stream",
1387
+ "text": [
1388
+ "/home/skiredj.abderrahman/.conda/envs/gptoss-vllm/lib/python3.10/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html\n",
1389
+ " from .autonotebook import tqdm as notebook_tqdm\n",
1390
+ "`torch_dtype` is deprecated! Use `dtype` instead!\n"
1391
+ ]
1392
+ },
1393
+ {
1394
+ "name": "stdout",
1395
+ "output_type": "stream",
1396
+ "text": [
1397
+ "Loading model...\n"
1398
+ ]
1399
+ },
1400
+ {
1401
+ "name": "stderr",
1402
+ "output_type": "stream",
1403
+ "text": [
1404
+ "Loading checkpoint shards: 100%|██████████| 2/2 [00:05<00:00, 2.92s/it]\n"
1405
+ ]
1406
+ },
1407
+ {
1408
+ "name": "stdout",
1409
+ "output_type": "stream",
1410
+ "text": [
1411
+ "Model loaded successfully!\n",
1412
+ "\n"
1413
+ ]
1414
+ }
1415
+ ],
1416
+ "source": [
1417
+ "import os\n",
1418
+ "from pathlib import Path\n",
1419
+ "from PIL import Image\n",
1420
+ "from transformers import AutoTokenizer, AutoProcessor, AutoModelForImageTextToText\n",
1421
+ "from datetime import datetime\n",
1422
+ "\n",
1423
+ "# Configuration\n",
1424
+ "os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0'\n",
1425
+ "model_path = \"Nanonets-OCR2-3B\" # Update this with your model path\n",
1426
+ "input_folder = \"table_images\" # Folder containing images\n",
1427
+ "output_file = \"output/Nanonets-OCR2-3B_output_skirdje_prompt_tables.txt\"\n",
1428
+ "# Supported image formats\n",
1429
+ "SUPPORTED_FORMATS = {'.png', '.jpg', '.jpeg', '.bmp', '.tiff', '.tif', '.webp'}\n",
1430
+ "\n",
1431
+ "# Load model\n",
1432
+ "print(\"Loading model...\")\n",
1433
+ "model = AutoModelForImageTextToText.from_pretrained(\n",
1434
+ " model_path, \n",
1435
+ " torch_dtype=\"auto\", \n",
1436
+ " device_map=\"auto\",\n",
1437
+ " offload_buffers=True, # Add this line\n",
1438
+ ")\n",
1439
+ "model.eval()\n",
1440
+ "tokenizer = AutoTokenizer.from_pretrained(model_path)\n",
1441
+ "processor = AutoProcessor.from_pretrained(model_path)\n",
1442
+ "print(\"Model loaded successfully!\\n\")\n",
1443
+ "\n",
1444
+ "def ocr_page_with_nanonets_s(image_path, model, processor, max_new_tokens=4096):\n",
1445
+ " prompt = \"\"\"\n",
1446
+ "Extract all visible text from the document image, following these rules:\n",
1447
+ "\n",
1448
+ "0. *Exhaustivity*\n",
1449
+ "\n",
1450
+ " * OCR *all text faithfully*. Do not omit or alter any word or detail.\n",
1451
+ "\n",
1452
+ "1. *Line Breaks*\n",
1453
+ "\n",
1454
+ " * Preserve line breaks *only* when:\n",
1455
+ "\n",
1456
+ " * The sentence ends with a period, question mark, or exclamation mark followed by a newline.\n",
1457
+ " * A heading/title is separated from body text.\n",
1458
+ " * ALWAYS insert a line break after subtitles starting with these keywords: المادة / الفصل / الباب.\n",
1459
+ " Example:\n",
1460
+ "\n",
1461
+ " \n",
1462
+ " المادة الثانية\n",
1463
+ " الفصل 2\n",
1464
+ " الباب الثالث\n",
1465
+ " \n",
1466
+ " * In all other cases, merge text into continuous paragraphs for readability.\n",
1467
+ "\n",
1468
+ "2. *Tables*\n",
1469
+ "\n",
1470
+ " * If tables are present, reformat them into clean, valid *Markdown table syntax*.\n",
1471
+ "\n",
1472
+ "3. *Ignore Content*\n",
1473
+ "\n",
1474
+ " * Do *not* OCR:\n",
1475
+ "\n",
1476
+ " * Headers at the very top of the page (e.g., journal name like \"الجريدة الرسمية\", issue numbers, dates).\n",
1477
+ " * Footers or page numbers.\n",
1478
+ " * Any number located at the extreme bottom of the page.\n",
1479
+ "\n",
1480
+ "4. *Subtitle Marking*\n",
1481
+ "\n",
1482
+ " * Subtitles are *only* those beginning with: المادة / الفصل.\n",
1483
+ " * Prepend ## (or deeper levels if nested) before subtitles, respecting hierarchy:\n",
1484
+ "\n",
1485
+ " * ##الفصل\n",
1486
+ "\n",
1487
+ " * ###المادة\n",
1488
+ "\n",
1489
+ " * Always add a line break *after* each subtitle. So whenever you write like \"المادة الثانية\" add a linebreak after it.\n",
1490
+ " * Never use a single # heading.\n",
1491
+ "\n",
1492
+ "5. *Two-Column Pages*\n",
1493
+ "\n",
1494
+ " * If the page is laid out in *two vertical columns* (with a clear dividing line between right and left), always OCR the *right column first, then the **left column*.\n",
1495
+ " * Do *not* confuse tables with columns — this rule applies *only* to page layouts, not tables.\n",
1496
+ "\n",
1497
+ "6. *Keep all dots*\n",
1498
+ " * If you see many dots in the text you are OCRizing (like \"....\"), please keep them all. They have meanings, so do not remove them.\n",
1499
+ "\n",
1500
+ "\"\"\"\n",
1501
+ " \n",
1502
+ " image = Image.open(image_path)\n",
1503
+ " messages = [\n",
1504
+ " {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n",
1505
+ " {\"role\": \"user\", \"content\": [\n",
1506
+ " {\"type\": \"image\", \"image\": f\"file://{image_path}\"},\n",
1507
+ " {\"type\": \"text\", \"text\": prompt},\n",
1508
+ " ]},\n",
1509
+ " ]\n",
1510
+ " text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)\n",
1511
+ " inputs = processor(text=[text], images=[image], padding=True, return_tensors=\"pt\")\n",
1512
+ " inputs = inputs.to(model.device)\n",
1513
+ " \n",
1514
+ " output_ids = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)\n",
1515
+ " generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]\n",
1516
+ " \n",
1517
+ " output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)\n",
1518
+ " return output_text[0]\n",
1519
+ "\n",
1520
+ "def process_folder(input_folder, output_file, model, processor, max_new_tokens=15000):\n",
1521
+ " # Get all image files from folder\n",
1522
+ " image_files = []\n",
1523
+ " for ext in SUPPORTED_FORMATS:\n",
1524
+ " image_files.extend(Path(input_folder).glob(f\"*{ext}\"))\n",
1525
+ " image_files.extend(Path(input_folder).glob(f\"*{ext.upper()}\"))\n",
1526
+ " \n",
1527
+ " # Sort files by name\n",
1528
+ " image_files = sorted(image_files)\n",
1529
+ " \n",
1530
+ " if not image_files:\n",
1531
+ " print(f\"No images found in {input_folder}\")\n",
1532
+ " return\n",
1533
+ " \n",
1534
+ " print(f\"Found {len(image_files)} images to process\\n\")\n",
1535
+ " \n",
1536
+ " # Process each image and write to output file\n",
1537
+ " with open(output_file, 'w', encoding='utf-8') as f:\n",
1538
+ " f.write(f\"OCR Results - Generated on {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\\n\")\n",
1539
+ " f.write(\"=\" * 80 + \"\\n\\n\")\n",
1540
+ " \n",
1541
+ " for idx, image_path in enumerate(image_files, 1):\n",
1542
+ " print(f\"Processing [{idx}/{len(image_files)}]: {image_path.name}\")\n",
1543
+ " \n",
1544
+ " try:\n",
1545
+ " result = ocr_page_with_nanonets_s(str(image_path), model, processor, max_new_tokens)\n",
1546
+ " \n",
1547
+ " # Write to file\n",
1548
+ " f.write(f\"{'=' * 80}\\n\")\n",
1549
+ " f.write(f\"File: {image_path.name}\\n\")\n",
1550
+ " f.write(f\"{'=' * 80}\\n\\n\")\n",
1551
+ " f.write(result)\n",
1552
+ " f.write(f\"\\n\\n{'=' * 80}\\n\\n\")\n",
1553
+ " \n",
1554
+ " print(f\"✓ Completed: {image_path.name}\\n\")\n",
1555
+ " \n",
1556
+ " except Exception as e:\n",
1557
+ " error_msg = f\"✗ Error processing {image_path.name}: {str(e)}\"\n",
1558
+ " print(error_msg)\n",
1559
+ " f.write(f\"\\n{error_msg}\\n\\n\")\n",
1560
+ " \n",
1561
+ " print(f\"\\n{'=' * 80}\")\n",
1562
+ " print(f\"Processing complete! Results saved to: {output_file}\")\n",
1563
+ " print(f\"{'=' * 80}\")\n"
1564
+ ]
1565
+ },
1566
+ {
1567
+ "cell_type": "code",
1568
+ "execution_count": 2,
1569
+ "id": "4edcf0f0-17a5-43e2-94a5-9d6ba7ebe5d7",
1570
+ "metadata": {},
1571
+ "outputs": [
1572
+ {
1573
+ "name": "stderr",
1574
+ "output_type": "stream",
1575
+ "text": [
1576
+ "The following generation flags are not valid and may be ignored: ['temperature']. Set `TRANSFORMERS_VERBOSITY=info` for more details.\n"
1577
+ ]
1578
+ },
1579
+ {
1580
+ "name": "stdout",
1581
+ "output_type": "stream",
1582
+ "text": [
1583
+ "Found 10 images to process\n",
1584
+ "\n",
1585
+ "Processing [1/10]: 1.PNG\n",
1586
+ "✓ Completed: 1.PNG\n",
1587
+ "\n",
1588
+ "Processing [2/10]: 10.PNG\n",
1589
+ "✓ Completed: 10.PNG\n",
1590
+ "\n",
1591
+ "Processing [3/10]: 2.PNG\n",
1592
+ "✓ Completed: 2.PNG\n",
1593
+ "\n",
1594
+ "Processing [4/10]: 3.PNG\n",
1595
+ "✓ Completed: 3.PNG\n",
1596
+ "\n",
1597
+ "Processing [5/10]: 4.PNG\n",
1598
+ "✓ Completed: 4.PNG\n",
1599
+ "\n",
1600
+ "Processing [6/10]: 5.PNG\n",
1601
+ "✓ Completed: 5.PNG\n",
1602
+ "\n",
1603
+ "Processing [7/10]: 6.PNG\n",
1604
+ "✓ Completed: 6.PNG\n",
1605
+ "\n",
1606
+ "Processing [8/10]: 7.PNG\n",
1607
+ "✓ Completed: 7.PNG\n",
1608
+ "\n",
1609
+ "Processing [9/10]: 8.PNG\n",
1610
+ "✓ Completed: 8.PNG\n",
1611
+ "\n",
1612
+ "Processing [10/10]: 9.PNG\n",
1613
+ "✓ Completed: 9.PNG\n",
1614
+ "\n",
1615
+ "\n",
1616
+ "================================================================================\n",
1617
+ "Processing complete! Results saved to: output/Nanonets-OCR2-3B_output_skirdje_prompt_tables.txt\n",
1618
+ "================================================================================\n"
1619
+ ]
1620
+ }
1621
+ ],
1622
+ "source": [
1623
+ "process_folder(input_folder, output_file, model, processor, max_new_tokens=15000)"
1624
+ ]
1625
+ },
1626
+ {
1627
+ "cell_type": "markdown",
1628
+ "id": "5fcb571c-8e83-4fdb-8f97-b2968d4b7007",
1629
+ "metadata": {},
1630
+ "source": [
1631
+ "# Test 3 : PaddleOCR-VL"
1632
+ ]
1633
+ },
1634
+ {
1635
+ "cell_type": "markdown",
1636
+ "id": "aed80449-5f57-460c-90ae-9a52156eeb15",
1637
+ "metadata": {},
1638
+ "source": [
1639
+ "## one single image code"
1640
+ ]
1641
+ },
1642
+ {
1643
+ "cell_type": "code",
1644
+ "execution_count": 19,
1645
+ "id": "8115ccb0-34b3-4bc4-a787-02970ca3d7d8",
1646
+ "metadata": {},
1647
+ "outputs": [],
1648
+ "source": [
1649
+ "from PIL import Image\n",
1650
+ "import torch\n",
1651
+ "from transformers import AutoModelForCausalLM, AutoProcessor\n"
1652
+ ]
1653
+ },
1654
+ {
1655
+ "cell_type": "code",
1656
+ "execution_count": 20,
1657
+ "id": "1563b7fc-1079-45b4-8c32-fc8708f28be9",
1658
+ "metadata": {},
1659
+ "outputs": [],
1660
+ "source": [
1661
+ "CHOSEN_TASK = \"ocr\" # Options: 'ocr' | 'table' | 'chart' | 'formula'\n",
1662
+ "PROMPTS = {\n",
1663
+ " \"ocr\": \"OCR:\",\n",
1664
+ " \"table\": \"Table Recognition:\",\n",
1665
+ " \"formula\": \"Formula Recognition:\",\n",
1666
+ " \"chart\": \"Chart Recognition:\",\n",
1667
+ "}"
1668
+ ]
1669
+ },
1670
+ {
1671
+ "cell_type": "code",
1672
+ "execution_count": 21,
1673
+ "id": "cbd40f39-60cd-4aa3-bbd8-0437e240a200",
1674
+ "metadata": {},
1675
+ "outputs": [],
1676
+ "source": [
1677
+ "DEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n",
1678
+ "model_path = \"./PaddleOCR-VL\"\n",
1679
+ "image_path = \"journal_25_right.png\"\n",
1680
+ "image = Image.open(image_path).convert(\"RGB\")"
1681
+ ]
1682
+ },
1683
+ {
1684
+ "cell_type": "code",
1685
+ "execution_count": 22,
1686
+ "id": "f06179ef-e124-499c-9278-546d94c9d944",
1687
+ "metadata": {},
1688
+ "outputs": [],
1689
+ "source": [
1690
+ "model = AutoModelForCausalLM.from_pretrained(\n",
1691
+ " model_path, trust_remote_code=True, torch_dtype=torch.bfloat16\n",
1692
+ ").to(DEVICE).eval()\n",
1693
+ "processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)"
1694
+ ]
1695
+ },
1696
+ {
1697
+ "cell_type": "code",
1698
+ "execution_count": 23,
1699
+ "id": "5c5489c9-fa90-44fc-a61d-2e29944435bd",
1700
+ "metadata": {},
1701
+ "outputs": [],
1702
+ "source": [
1703
+ "messages = [\n",
1704
+ " {\"role\": \"user\", \n",
1705
+ " \"content\": [\n",
1706
+ " {\"type\": \"image\", \"image\": image},\n",
1707
+ " {\"type\": \"text\", \"text\": PROMPTS[CHOSEN_TASK]},\n",
1708
+ " ]\n",
1709
+ " }\n",
1710
+ "]\n",
1711
+ "inputs = processor.apply_chat_template(\n",
1712
+ " messages, \n",
1713
+ " tokenize=True, \n",
1714
+ " add_generation_prompt=True, \t\n",
1715
+ " return_dict=True,\n",
1716
+ " return_tensors=\"pt\"\n",
1717
+ ").to(DEVICE)"
1718
+ ]
1719
+ },
1720
+ {
1721
+ "cell_type": "code",
1722
+ "execution_count": 24,
1723
+ "id": "28ee94a9-102b-49d6-8aa3-dacb0e0ae878",
1724
+ "metadata": {},
1725
+ "outputs": [
1726
+ {
1727
+ "name": "stderr",
1728
+ "output_type": "stream",
1729
+ "text": [
1730
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1731
+ ]
1732
+ },
1733
+ {
1734
+ "name": "stdout",
1735
+ "output_type": "stream",
1736
+ "text": [
1737
+ "User: OCR:\n",
1738
+ "Assistant: الجريدة\n",
1739
+ "قرر ما يلي :\n",
1740
+ "المادة الأولى\n",
1741
+ "تقبل لمعادلة دبلوم مهندس دولة، الشهادة التالية في\n",
1742
+ "Génie civil\n",
1743
+ "-Titre d'ingénieur diplôme de l'Ecole nationale supérieure\n",
1744
+ "Mines-Télécom Lille Douai, de l'Institut Mines-Télécom, délivré en date du 10 janvier 2024 - France.\n",
1745
+ "المادة الثانية\n",
1746
+ "ينشر هذا القرار بالجريدة الرسمية.\n",
1747
+ "وحرر بالرباط في 19 من محرم 1446 يوليو 25 (2024).\n",
1748
+ "الإمضاء: عبد اللطيف ميراوي.\n",
1749
+ "قرار لوزير التعليم العالي والبحث العلمي والابتكار رقم 2061.24\n",
1750
+ "صادر في 19 من محرم 1446 يوليو 25 (2024) بتحديد بعض\n",
1751
+ "المعادلات بين الشهادات.\n",
1752
+ "وزير التعليم العالي والبحث العلمي والابتكار،\n",
1753
+ "بناء على المرسوم رقم 2.01.333 2.01.333 الصادر في 28 من ربيع الأول 1422\n",
1754
+ "بينو 21\n",
1755
+ "(2001\n",
1756
+ "معادلة شهادات التعليم العالي :\n",
1757
+ "وعلى المرسوم رقم 2.04.89\n",
1758
+ "(2004\n",
1759
+ "وتتميمه :\n",
1760
+ "وعلى المرسوم رقم 2.21.838\n",
1761
+ "(2021\n",
1762
+ "العلي والابتكار :\n",
1763
+ "وبعد استشارة اللجنة القطاعية للعلوم والتقنيات والهندسة\n",
1764
+ "والهندسة المعمارية المنعقدة بتاريخ 4 يوليو 2024\n",
1765
+ "قرر ما يلي :\n",
1766
+ "المادة الأولى\n",
1767
+ ":\n",
1768
+ "Informatique دبلوم مهندس دولة، الشهادة التالية في\n",
1769
+ "-Titre d'ingénieur de télécom Nancy de l'Université de Lorraine, délivré en date du 26 octobre 2023 - France.\n"
1770
+ ]
1771
+ }
1772
+ ],
1773
+ "source": [
1774
+ "outputs = model.generate(**inputs, max_new_tokens=1024)\n",
1775
+ "outputs = proceqstatsor.batch_decode(outputs, skip_special_tokens=True)[0]\n",
1776
+ "print(outputs)"
1777
+ ]
1778
+ },
1779
+ {
1780
+ "cell_type": "markdown",
1781
+ "id": "8de14b64-c171-4940-99af-d892443ea636",
1782
+ "metadata": {},
1783
+ "source": [
1784
+ "## for multiple images "
1785
+ ]
1786
+ },
1787
+ {
1788
+ "cell_type": "code",
1789
+ "execution_count": 16,
1790
+ "id": "5bcf22eb-f622-4d45-9efb-bdd2f75c413f",
1791
+ "metadata": {},
1792
+ "outputs": [
1793
+ {
1794
+ "name": "stdout",
1795
+ "output_type": "stream",
1796
+ "text": [
1797
+ "Loading model...\n",
1798
+ "Model loaded successfully!\n"
1799
+ ]
1800
+ }
1801
+ ],
1802
+ "source": [
1803
+ "from PIL import Image\n",
1804
+ "import torch\n",
1805
+ "from transformers import AutoModelForCausalLM, AutoProcessor\n",
1806
+ "import os\n",
1807
+ "from pathlib import Path\n",
1808
+ "\n",
1809
+ "CHOSEN_TASK = \"ocr\" # Options: 'ocr' | 'table' | 'chart' | 'formula'\n",
1810
+ "PROMPTS = {\n",
1811
+ " \"ocr\": \"OCR:\",\n",
1812
+ " \"table\": \"Table Recognition:\",\n",
1813
+ " \"formula\": \"Formula Recognition:\",\n",
1814
+ " \"chart\": \"Chart Recognition:\",\n",
1815
+ "}\n",
1816
+ "\n",
1817
+ "DEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n",
1818
+ "model_path = \"./PaddleOCR-VL\"\n",
1819
+ "images_folder = \"./images\"\n",
1820
+ "output_folder = \"./output\"\n",
1821
+ "output_file = os.path.join(output_folder, \"output_PaddleOCR-VL.txt\")\n",
1822
+ "\n",
1823
+ "# Create output folder if it doesn't exist\n",
1824
+ "os.makedirs(output_folder, exist_ok=True)\n",
1825
+ "\n",
1826
+ "# Load model once\n",
1827
+ "print(\"Loading model...\")\n",
1828
+ "model = AutoModelForCausalLM.from_pretrained(\n",
1829
+ " model_path, trust_remote_code=True, torch_dtype=torch.bfloat16\n",
1830
+ ").to(DEVICE).eval()\n",
1831
+ "processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)\n",
1832
+ "print(\"Model loaded successfully!\")"
1833
+ ]
1834
+ },
1835
+ {
1836
+ "cell_type": "code",
1837
+ "execution_count": 17,
1838
+ "id": "3069509a-01e6-453b-9ca0-3be757db1dd0",
1839
+ "metadata": {
1840
+ "scrolled": true
1841
+ },
1842
+ "outputs": [
1843
+ {
1844
+ "name": "stdout",
1845
+ "output_type": "stream",
1846
+ "text": [
1847
+ "Found 32 images to process\n",
1848
+ "Processing [1/32]: journal_10.png\n"
1849
+ ]
1850
+ },
1851
+ {
1852
+ "name": "stderr",
1853
+ "output_type": "stream",
1854
+ "text": [
1855
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1856
+ ]
1857
+ },
1858
+ {
1859
+ "name": "stdout",
1860
+ "output_type": "stream",
1861
+ "text": [
1862
+ " ✓ Completed: journal_10.png\n",
1863
+ "Processing [2/32]: journal_12.png\n"
1864
+ ]
1865
+ },
1866
+ {
1867
+ "name": "stderr",
1868
+ "output_type": "stream",
1869
+ "text": [
1870
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1871
+ ]
1872
+ },
1873
+ {
1874
+ "name": "stdout",
1875
+ "output_type": "stream",
1876
+ "text": [
1877
+ " ✓ Completed: journal_12.png\n",
1878
+ "Processing [3/32]: journal_13.png\n"
1879
+ ]
1880
+ },
1881
+ {
1882
+ "name": "stderr",
1883
+ "output_type": "stream",
1884
+ "text": [
1885
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1886
+ ]
1887
+ },
1888
+ {
1889
+ "name": "stdout",
1890
+ "output_type": "stream",
1891
+ "text": [
1892
+ " ✓ Completed: journal_13.png\n",
1893
+ "Processing [4/32]: journal_14.png\n"
1894
+ ]
1895
+ },
1896
+ {
1897
+ "name": "stderr",
1898
+ "output_type": "stream",
1899
+ "text": [
1900
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1901
+ ]
1902
+ },
1903
+ {
1904
+ "name": "stdout",
1905
+ "output_type": "stream",
1906
+ "text": [
1907
+ " ✓ Completed: journal_14.png\n",
1908
+ "Processing [5/32]: journal_15_left.png\n"
1909
+ ]
1910
+ },
1911
+ {
1912
+ "name": "stderr",
1913
+ "output_type": "stream",
1914
+ "text": [
1915
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1916
+ ]
1917
+ },
1918
+ {
1919
+ "name": "stdout",
1920
+ "output_type": "stream",
1921
+ "text": [
1922
+ " ✓ Completed: journal_15_left.png\n",
1923
+ "Processing [6/32]: journal_15_right.png\n"
1924
+ ]
1925
+ },
1926
+ {
1927
+ "name": "stderr",
1928
+ "output_type": "stream",
1929
+ "text": [
1930
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1931
+ ]
1932
+ },
1933
+ {
1934
+ "name": "stdout",
1935
+ "output_type": "stream",
1936
+ "text": [
1937
+ " ✓ Completed: journal_15_right.png\n",
1938
+ "Processing [7/32]: journal_16_left.png\n"
1939
+ ]
1940
+ },
1941
+ {
1942
+ "name": "stderr",
1943
+ "output_type": "stream",
1944
+ "text": [
1945
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1946
+ ]
1947
+ },
1948
+ {
1949
+ "name": "stdout",
1950
+ "output_type": "stream",
1951
+ "text": [
1952
+ " ✓ Completed: journal_16_left.png\n",
1953
+ "Processing [8/32]: journal_16_right.png\n"
1954
+ ]
1955
+ },
1956
+ {
1957
+ "name": "stderr",
1958
+ "output_type": "stream",
1959
+ "text": [
1960
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1961
+ ]
1962
+ },
1963
+ {
1964
+ "name": "stdout",
1965
+ "output_type": "stream",
1966
+ "text": [
1967
+ " ✓ Completed: journal_16_right.png\n",
1968
+ "Processing [9/32]: journal_17_left.png\n"
1969
+ ]
1970
+ },
1971
+ {
1972
+ "name": "stderr",
1973
+ "output_type": "stream",
1974
+ "text": [
1975
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1976
+ ]
1977
+ },
1978
+ {
1979
+ "name": "stdout",
1980
+ "output_type": "stream",
1981
+ "text": [
1982
+ " ✓ Completed: journal_17_left.png\n",
1983
+ "Processing [10/32]: journal_17_right.png\n"
1984
+ ]
1985
+ },
1986
+ {
1987
+ "name": "stderr",
1988
+ "output_type": "stream",
1989
+ "text": [
1990
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
1991
+ ]
1992
+ },
1993
+ {
1994
+ "name": "stdout",
1995
+ "output_type": "stream",
1996
+ "text": [
1997
+ " ✓ Completed: journal_17_right.png\n",
1998
+ "Processing [11/32]: journal_18_left.png\n"
1999
+ ]
2000
+ },
2001
+ {
2002
+ "name": "stderr",
2003
+ "output_type": "stream",
2004
+ "text": [
2005
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2006
+ ]
2007
+ },
2008
+ {
2009
+ "name": "stdout",
2010
+ "output_type": "stream",
2011
+ "text": [
2012
+ " ✓ Completed: journal_18_left.png\n",
2013
+ "Processing [12/32]: journal_18_right.png\n"
2014
+ ]
2015
+ },
2016
+ {
2017
+ "name": "stderr",
2018
+ "output_type": "stream",
2019
+ "text": [
2020
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2021
+ ]
2022
+ },
2023
+ {
2024
+ "name": "stdout",
2025
+ "output_type": "stream",
2026
+ "text": [
2027
+ " ✓ Completed: journal_18_right.png\n",
2028
+ "Processing [13/32]: journal_19.png\n"
2029
+ ]
2030
+ },
2031
+ {
2032
+ "name": "stderr",
2033
+ "output_type": "stream",
2034
+ "text": [
2035
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2036
+ ]
2037
+ },
2038
+ {
2039
+ "name": "stdout",
2040
+ "output_type": "stream",
2041
+ "text": [
2042
+ " ✓ Completed: journal_19.png\n",
2043
+ "Processing [14/32]: journal_1_left.png\n"
2044
+ ]
2045
+ },
2046
+ {
2047
+ "name": "stderr",
2048
+ "output_type": "stream",
2049
+ "text": [
2050
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2051
+ ]
2052
+ },
2053
+ {
2054
+ "name": "stdout",
2055
+ "output_type": "stream",
2056
+ "text": [
2057
+ " ✓ Completed: journal_1_left.png\n",
2058
+ "Processing [15/32]: journal_1_right.png\n"
2059
+ ]
2060
+ },
2061
+ {
2062
+ "name": "stderr",
2063
+ "output_type": "stream",
2064
+ "text": [
2065
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2066
+ ]
2067
+ },
2068
+ {
2069
+ "name": "stdout",
2070
+ "output_type": "stream",
2071
+ "text": [
2072
+ " ✓ Completed: journal_1_right.png\n",
2073
+ "Processing [16/32]: journal_2.png\n"
2074
+ ]
2075
+ },
2076
+ {
2077
+ "name": "stderr",
2078
+ "output_type": "stream",
2079
+ "text": [
2080
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2081
+ ]
2082
+ },
2083
+ {
2084
+ "name": "stdout",
2085
+ "output_type": "stream",
2086
+ "text": [
2087
+ " ✓ Completed: journal_2.png\n",
2088
+ "Processing [17/32]: journal_20.png\n"
2089
+ ]
2090
+ },
2091
+ {
2092
+ "name": "stderr",
2093
+ "output_type": "stream",
2094
+ "text": [
2095
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2096
+ ]
2097
+ },
2098
+ {
2099
+ "name": "stdout",
2100
+ "output_type": "stream",
2101
+ "text": [
2102
+ " ✓ Completed: journal_20.png\n",
2103
+ "Processing [18/32]: journal_21.png\n"
2104
+ ]
2105
+ },
2106
+ {
2107
+ "name": "stderr",
2108
+ "output_type": "stream",
2109
+ "text": [
2110
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2111
+ ]
2112
+ },
2113
+ {
2114
+ "name": "stdout",
2115
+ "output_type": "stream",
2116
+ "text": [
2117
+ " ✓ Completed: journal_21.png\n",
2118
+ "Processing [19/32]: journal_22.png\n"
2119
+ ]
2120
+ },
2121
+ {
2122
+ "name": "stderr",
2123
+ "output_type": "stream",
2124
+ "text": [
2125
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2126
+ ]
2127
+ },
2128
+ {
2129
+ "name": "stdout",
2130
+ "output_type": "stream",
2131
+ "text": [
2132
+ " ✓ Completed: journal_22.png\n",
2133
+ "Processing [20/32]: journal_23.png\n"
2134
+ ]
2135
+ },
2136
+ {
2137
+ "name": "stderr",
2138
+ "output_type": "stream",
2139
+ "text": [
2140
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2141
+ ]
2142
+ },
2143
+ {
2144
+ "name": "stdout",
2145
+ "output_type": "stream",
2146
+ "text": [
2147
+ " ✓ Completed: journal_23.png\n",
2148
+ "Processing [21/32]: journal_24_left.png\n"
2149
+ ]
2150
+ },
2151
+ {
2152
+ "name": "stderr",
2153
+ "output_type": "stream",
2154
+ "text": [
2155
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2156
+ ]
2157
+ },
2158
+ {
2159
+ "name": "stdout",
2160
+ "output_type": "stream",
2161
+ "text": [
2162
+ " ✓ Completed: journal_24_left.png\n",
2163
+ "Processing [22/32]: journal_24_right.png\n"
2164
+ ]
2165
+ },
2166
+ {
2167
+ "name": "stderr",
2168
+ "output_type": "stream",
2169
+ "text": [
2170
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2171
+ ]
2172
+ },
2173
+ {
2174
+ "name": "stdout",
2175
+ "output_type": "stream",
2176
+ "text": [
2177
+ " ✓ Completed: journal_24_right.png\n",
2178
+ "Processing [23/32]: journal_25_left.png\n"
2179
+ ]
2180
+ },
2181
+ {
2182
+ "name": "stderr",
2183
+ "output_type": "stream",
2184
+ "text": [
2185
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2186
+ ]
2187
+ },
2188
+ {
2189
+ "name": "stdout",
2190
+ "output_type": "stream",
2191
+ "text": [
2192
+ " ✓ Completed: journal_25_left.png\n",
2193
+ "Processing [24/32]: journal_25_right.png\n"
2194
+ ]
2195
+ },
2196
+ {
2197
+ "name": "stderr",
2198
+ "output_type": "stream",
2199
+ "text": [
2200
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2201
+ ]
2202
+ },
2203
+ {
2204
+ "name": "stdout",
2205
+ "output_type": "stream",
2206
+ "text": [
2207
+ " ✓ Completed: journal_25_right.png\n",
2208
+ "Processing [25/32]: journal_27.png\n"
2209
+ ]
2210
+ },
2211
+ {
2212
+ "name": "stderr",
2213
+ "output_type": "stream",
2214
+ "text": [
2215
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2216
+ ]
2217
+ },
2218
+ {
2219
+ "name": "stdout",
2220
+ "output_type": "stream",
2221
+ "text": [
2222
+ " ✓ Completed: journal_27.png\n",
2223
+ "Processing [26/32]: journal_3.png\n"
2224
+ ]
2225
+ },
2226
+ {
2227
+ "name": "stderr",
2228
+ "output_type": "stream",
2229
+ "text": [
2230
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2231
+ ]
2232
+ },
2233
+ {
2234
+ "name": "stdout",
2235
+ "output_type": "stream",
2236
+ "text": [
2237
+ " ✓ Completed: journal_3.png\n",
2238
+ "Processing [27/32]: journal_4.png\n"
2239
+ ]
2240
+ },
2241
+ {
2242
+ "name": "stderr",
2243
+ "output_type": "stream",
2244
+ "text": [
2245
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2246
+ ]
2247
+ },
2248
+ {
2249
+ "name": "stdout",
2250
+ "output_type": "stream",
2251
+ "text": [
2252
+ " ✓ Completed: journal_4.png\n",
2253
+ "Processing [28/32]: journal_5.png\n"
2254
+ ]
2255
+ },
2256
+ {
2257
+ "name": "stderr",
2258
+ "output_type": "stream",
2259
+ "text": [
2260
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2261
+ ]
2262
+ },
2263
+ {
2264
+ "name": "stdout",
2265
+ "output_type": "stream",
2266
+ "text": [
2267
+ " ✓ Completed: journal_5.png\n",
2268
+ "Processing [29/32]: journal_6.png\n"
2269
+ ]
2270
+ },
2271
+ {
2272
+ "name": "stderr",
2273
+ "output_type": "stream",
2274
+ "text": [
2275
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2276
+ ]
2277
+ },
2278
+ {
2279
+ "name": "stdout",
2280
+ "output_type": "stream",
2281
+ "text": [
2282
+ " ✓ Completed: journal_6.png\n",
2283
+ "Processing [30/32]: journal_7.png\n"
2284
+ ]
2285
+ },
2286
+ {
2287
+ "name": "stderr",
2288
+ "output_type": "stream",
2289
+ "text": [
2290
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2291
+ ]
2292
+ },
2293
+ {
2294
+ "name": "stdout",
2295
+ "output_type": "stream",
2296
+ "text": [
2297
+ " ✓ Completed: journal_7.png\n",
2298
+ "Processing [31/32]: journal_8.png\n"
2299
+ ]
2300
+ },
2301
+ {
2302
+ "name": "stderr",
2303
+ "output_type": "stream",
2304
+ "text": [
2305
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2306
+ ]
2307
+ },
2308
+ {
2309
+ "name": "stdout",
2310
+ "output_type": "stream",
2311
+ "text": [
2312
+ " ✓ Completed: journal_8.png\n",
2313
+ "Processing [32/32]: journal_9.png\n"
2314
+ ]
2315
+ },
2316
+ {
2317
+ "name": "stderr",
2318
+ "output_type": "stream",
2319
+ "text": [
2320
+ "Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.\n"
2321
+ ]
2322
+ },
2323
+ {
2324
+ "name": "stdout",
2325
+ "output_type": "stream",
2326
+ "text": [
2327
+ " ✓ Completed: journal_9.png\n",
2328
+ "\n",
2329
+ "✓ All images processed!\n",
2330
+ "Output saved to: ./output/output_PaddleOCR-VL.txt\n"
2331
+ ]
2332
+ }
2333
+ ],
2334
+ "source": [
2335
+ "# Get all image files from the folder\n",
2336
+ "image_extensions = {'.png', '.jpg', '.jpeg', '.bmp', '.gif', '.tiff', '.webp'}\n",
2337
+ "image_files = [\n",
2338
+ " f for f in os.listdir(images_folder) \n",
2339
+ " if os.path.splitext(f.lower())[1] in image_extensions\n",
2340
+ "]\n",
2341
+ "image_files.sort() # Sort files for consistent ordering\n",
2342
+ "\n",
2343
+ "print(f\"Found {len(image_files)} images to process\")\n",
2344
+ "\n",
2345
+ "# Open output file for writing\n",
2346
+ "with open(output_file, 'w', encoding='utf-8') as f:\n",
2347
+ " for idx, image_file in enumerate(image_files, 1):\n",
2348
+ " image_path = os.path.join(images_folder, image_file)\n",
2349
+ " print(f\"Processing [{idx}/{len(image_files)}]: {image_file}\")\n",
2350
+ " \n",
2351
+ " try:\n",
2352
+ " # Load and process image\n",
2353
+ " image = Image.open(image_path).convert(\"RGB\")\n",
2354
+ " \n",
2355
+ " messages = [\n",
2356
+ " {\n",
2357
+ " \"role\": \"user\",\n",
2358
+ " \"content\": [\n",
2359
+ " {\"type\": \"image\", \"image\": image},\n",
2360
+ " {\"type\": \"text\", \"text\": PROMPTS[CHOSEN_TASK]},\n",
2361
+ " ]\n",
2362
+ " }\n",
2363
+ " ]\n",
2364
+ " \n",
2365
+ " inputs = processor.apply_chat_template(\n",
2366
+ " messages,\n",
2367
+ " tokenize=True,\n",
2368
+ " add_generation_prompt=True,\n",
2369
+ " return_dict=True,\n",
2370
+ " return_tensors=\"pt\"\n",
2371
+ " ).to(DEVICE)\n",
2372
+ " \n",
2373
+ " outputs = model.generate(**inputs, max_new_tokens=1024)\n",
2374
+ " result = processor.batch_decode(outputs, skip_special_tokens=True)[0]\n",
2375
+ " \n",
2376
+ " # Write to file with separator\n",
2377
+ " f.write(f\"{'='*80}\\n\")\n",
2378
+ " f.write(f\"Image: {image_file}\\n\")\n",
2379
+ " f.write(f\"{'='*80}\\n\")\n",
2380
+ " f.write(result)\n",
2381
+ " f.write(f\"\\n\\n\")\n",
2382
+ " \n",
2383
+ " print(f\" ✓ Completed: {image_file}\")\n",
2384
+ " \n",
2385
+ " except Exception as e:\n",
2386
+ " error_msg = f\"Error processing {image_file}: {str(e)}\"\n",
2387
+ " print(f\" ✗ {error_msg}\")\n",
2388
+ " f.write(f\"{'='*80}\\n\")\n",
2389
+ " f.write(f\"Image: {image_file}\\n\")\n",
2390
+ " f.write(f\"{'='*80}\\n\")\n",
2391
+ " f.write(f\"ERROR: {str(e)}\\n\\n\")\n",
2392
+ "\n",
2393
+ "print(f\"\\n✓ All images processed!\")\n",
2394
+ "print(f\"Output saved to: {output_file}\")"
2395
+ ]
2396
+ }
2397
+ ],
2398
+ "metadata": {
2399
+ "kernelspec": {
2400
+ "display_name": "Python 3 (ipykernel)",
2401
+ "language": "python",
2402
+ "name": "python3"
2403
+ },
2404
+ "language_info": {
2405
+ "codemirror_mode": {
2406
+ "name": "ipython",
2407
+ "version": 3
2408
+ },
2409
+ "file_extension": ".py",
2410
+ "mimetype": "text/x-python",
2411
+ "name": "python",
2412
+ "nbconvert_exporter": "python",
2413
+ "pygments_lexer": "ipython3",
2414
+ "version": "3.10.19"
2415
+ }
2416
+ },
2417
+ "nbformat": 4,
2418
+ "nbformat_minor": 5
2419
+ }