uzocr-8b: OCR for scanned Uzbek documents

uzocr-8b reads scanned pages of Uzbek books, dissertations and dissertation abstracts and returns their text as Markdown.

  • It handles both Uzbek scripts, Cyrillic (including Ñž, Ò›, Ò“, Ò³) and Latin (oÊ», gÊ», ʼ), on typewritten and printed pages, including low-quality scans.
  • It preserves the page structure: headings, paragraphs, tables, footnotes, formulas (LaTeX) and the table of contents.
  • The model runs fully locally on a single GPU.

O'zbekcha. uzocr-8b — skanerlangan o'zbek kitoblari, dissertatsiya va avtoreferatlar betini o'qib, matnini Markdown ko'rinishida qaytaradigan model. U kirill va lotin yozuvidagi, mashinkada yoki bosmaxonada terilgan, sifati past skanlar bilan ham ishlaydi. Sarlavha, jadval, snoska, formula va mundarijani saqlaydi. Bitta GPU'da to'liq lokal ishlaydi.

Base model Qwen/Qwen3-VL-8B-Instruct
Method LoRA fine-tuning, merged into the weights (bf16)
Parameters 8B
Languages Uzbek (Cyrillic, Latin); Russian and English passages inside Uzbek documents
Output Markdown (tables, footnotes [^n], LaTeX formulas)
GGUF version JahongirB/uzocr-8b-GGUF (llama.cpp, LM Studio, Ollama)
License Apache-2.0

Results

Test set Pages CER, % WER, % Word accuracy, %
— — — — —

Results will be added.

Usage

The model was trained with the system prompt below. Use it verbatim, with greedy decoding (temperature=0).

System prompt
You are a high-precision OCR engine for scanned BOOK pages written in Uzbek (Latin or Cyrillic), sometimes mixed with Russian or English. Pages may contain printed or typewritten text, handwritten notes, tables, poems, footnotes, stamps and images.

TASK: Transcribe the page image EXACTLY as written and return clean Markdown.

RULES:
1. Verbatim only. Do NOT translate, summarize, paraphrase, modernize spelling, fix grammar or complete sentences. Cyrillic stays Cyrillic, Latin stays Latin. Keep Uzbek Cyrillic letters Ñž Ò› Ò“ Ò³ exactly as printed.
2. Uzbek Latin: write oʻ and gʻ with U+02BB (ʻ) and the tutuq belgisi with U+02BC (ʼ). Never ó, ğ or backticks.
3. Handwriting: if it is a margin or interlinear note added to printed text, write it as {qoʻlyozma: ...} where it appears.
4. Uncertain word → best reading + [?] (Toshkent[?]). Unreadable → [oʻqib boʻlmaydi]. Never guess silently.
5. Book structure:
   - Chapter titles → "# ", section titles → "## ". Only for real visual headings.
   - Join all lines of a paragraph into one line; separate paragraphs with a blank line. A word split by a hyphen at a line end is joined (kitob-/lar → kitoblar); real hyphenated words keep the hyphen.
   - Poems: keep every line break; separate stanzas with a blank line.
   - Footnotes at the bottom → write at the end as [^n]: text, and keep the marker [^n] in the body where the number appears.
   - Tables → Markdown tables with all rows and columns.
   - Table of contents: write "title — page"; never reproduce dot leaders.
   - Drop caps (large first letter) → merge into the word.
   - Images/illustrations → [rasm: short caption if printed].
6. Only if the page really has a running header/footer or page number, put it on a separate first/last line prefixed with "HEADER:" / "FOOTER:". If there is none, do not write these lines.
7. Output ONLY the transcription. No comments, no code fences.

User message: the page image plus one of these texts:

  • Cyrillic: Transcribe this page. This book is printed in Uzbek CYRILLIC script: keep every word in Cyrillic exactly as printed; never transliterate to Latin.
  • Latin: Transcribe this page. This book is printed in Uzbek LATIN script: keep every word in Latin exactly as printed; never transliterate to Cyrillic.

Render pages at about 200 dpi (up to about 2.6–2.9 megapixels).

vLLM (recommended)

vllm serve JahongirB/uzocr-8b --served-model-name uzocr-8b --dtype bfloat16 --max-model-len 8192 \
  --max-num-seqs 24 --limit-mm-per-prompt.image 1 \
  --mm-processor-kwargs.min_pixels 200704 --mm-processor-kwargs.max_pixels 2600000 \
  --generation-config vllm
import base64, openai

SYSTEM_PROMPT = "..."  # the system prompt above
client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="-")
img = base64.b64encode(open("page.png", "rb").read()).decode()
r = client.chat.completions.create(
    model="uzocr-8b", temperature=0, max_tokens=4096,
    messages=[{"role": "system", "content": SYSTEM_PROMPT},
              {"role": "user", "content": [
                  {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img}"}},
                  {"type": "text", "text": "Transcribe this page. This book is printed in Uzbek CYRILLIC script: keep every word in Cyrillic exactly as printed; never transliterate to Latin."}]}])
print(r.choices[0].message.content)

llama.cpp (12 GB GPU)

Use uzocr-8b-GGUF:

llama-server -m uzocr-8b-v3b-Q8_0.gguf --mmproj mmproj-uzocr-8b-v3-F16.gguf -c 8192 -np 4 --temp 0

Requirements

  • vLLM, bf16: about 17 GB for the weights plus the KV cache. With 24 concurrent requests, peak VRAM was 30 GB on an RTX 5090.
  • llama.cpp, Q8_0: 8.7 GB for the model plus 1.2 GB for the vision projector.

Training

  • Data: page images of Uzbek documents (Cyrillic and Latin, typewritten and printed, scans of varying quality) paired with Markdown transcriptions in the format of the system prompt above. Documents used for evaluation were excluded from training; this was checked with text-shingle overlap.
  • Method: LoRA, r = 32, alpha = 64, dropout 0.05, on all attention and MLP projections of the language model. The vision encoder was frozen.
  • Optimization: loss on answer tokens only, cosine schedule, gradient accumulation 8. Three stages (lr 1e-4, then 5e-5 and 5e-5). This release is the last checkpoint of stage 3.
  • Compute: about 17.6 GPU-hours on a single NVIDIA RTX 5090.

Limitations

  • Most remaining errors involve the Uzbek-specific letters Ñž/Ò›/Ò“/Ò³ (e.g. Ò› → к).
  • Pseudo-graphic typewriter tables can cause repetition loops. A loop check and a retry are recommended.
  • Pages that mix scripts (e.g. an English summary page in a Cyrillic book) may be transliterated into the main script.
  • Handwritten text is not a target of the model.
  • Russian pages are readable, but the model was not tuned for them.

License

Apache-2.0, the same as the base model.

Downloads last month
54
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JahongirB/uzocr-8b

Adapter
(220)
this model
Adapters
1 model
Quantizations
1 model