---
license: agpl-3.0
base_model: Qwen/Qwen3.5-4B
base_model_relation: finetune
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- ocr
- document-parsing
- document-understanding
- markdown
- latex
- table-recognition
- multilingual
- vllm
- qwen3_5
language:
- en
- de
- es
- fr
- id
- it
- nl
- pt
- vi
- ar
- hi
- ja
- ko
- ru
- th
- zh
---
# PepperOCR-VL
PepperOCR-VL turns a page image into clean Markdown. Point it at a scanned or
photographed document, an invoice, a paper, a textbook page or a slide, and it
returns the text in reading order with headings, lists, tables and formulas
preserved. It is a 4.5B-parameter vision-language model fine-tuned from
[Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) for 17 languages across
Latin and non-Latin scripts. On [MDPBench](https://github.com/Yuliang-Liu/MultimodalOCR/tree/main/MDPBench)
it holds the **state-of-the-art Korean (92.6) and Thai (83.3) scores** across
all listed models, including large general-purpose VLMs, as of the
September 2026 leaderboard.
## About Sionic AI
PepperOCR-VL is built by [Sionic AI](https://sionic.ai), a Seoul-based AI
company building document and language models for enterprise use. If you are
building agentic document parsing, large-scale document processing pipelines,
or want to use PepperOCR in a commercial product, we would like to hear from
you: contact us through the website. Commercial licensing, larger deployments
and hosted inference are available.
## What it outputs
One page image in, one Markdown document out:
- **Text and layout**: headings (`#`), paragraphs, bullet and numbered lists,
in the original language and reading order. No translation, no summarizing,
no spelling correction.
- **Tables**: HTML `` blocks with `rowspan`/`colspan`, which survive
merged cells better than Markdown pipe tables.
- **Mathematics**: LaTeX, `$x$` for inline formulas and `$$x$$` for display
formulas.
- Nothing else: no preamble, no explanation, just the transcription.
## Prompt format
The model is trained to respond to a single user turn that contains the page
image followed by this instruction. Use it verbatim (also in
`prompt.txt`):
```text
You are an advanced hybrid OCR engine capable of processing multilingual text mixed with mathematical notation. Your goal is to transcribe the content with high fidelity.Strict Rules: 1. Multilingual Precision: Transcribe text exactly as it appears in the original language. Do not translate, summarize, or correct original spelling errors. 2. Math Formatting: Identify all mathematical expressions and convert them into LaTeX. 3. Use single dollar signs ($x$) for inline math (formulas within a sentence). 4. Use double dollar signs ($$x$$) for display math (standalone formulas on their own lines). 5. Layout & Structure: Use Markdown to preserve the visual structure (headers, paragraphs, lists). 6. Output Only: Output the transcribed text directly without any conversational filler.
```
Message layout:
```json
{"role": "user", "content": [
{"type": "image", "image": ""},
{"type": "text", "text": ""}
]}
```
Keep thinking mode off (`enable_thinking: false`); the model was tuned and
evaluated without it. The decoding settings are in the next section.
## Inference settings
The decoding defaults that produced the reported scores are built into
`generation_config.json` (greedy decoding, `repetition_penalty` 1.05,
`no_repeat_ngram_size` 30, up to 8192 new tokens), so `transformers` and vLLM
apply them automatically. Keep thinking mode off.
The complete configuration behind the benchmark numbers, including the vLLM
serving arguments, image preprocessing, the orientation threshold, the
runaway-output retry policy and the exact software versions, is in
[`inference_config.yaml`](inference_config.yaml). Start from that file if you
want to reproduce the scores or tune the setup.
## Quickstart
### vLLM
```bash
vllm serve sionic-ai/PepperOCR-VL \
--served-model-name pepperocr-vl \
--max-model-len 32768 \
--max-num-seqs 64 \
--gpu-memory-utilization 0.85 \
--gdn-prefill-backend triton # vLLM 0.24; drop if your version does not have it
```
```python
import base64, pathlib
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
prompt = pathlib.Path("prompt.txt").read_text()
image = base64.b64encode(pathlib.Path("page.jpg").read_bytes()).decode()
response = client.chat.completions.create(
model="pepperocr-vl",
messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
{"type": "text", "text": prompt},
]}],
temperature=0.0,
top_p=1.0,
max_tokens=8192,
seed=42,
presence_penalty=0.0,
extra_body={
"top_k": -1,
"repetition_penalty": 1.05,
"no_repeat_ngram_size": 30,
"chat_template_kwargs": {"enable_thinking": False},
},
)
print(response.choices[0].message.content)
```
### Transformers
```python
import torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "sionic-ai/PepperOCR-VL"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="cuda"
)
prompt = open("prompt.txt").read()
messages = [{"role": "user", "content": [
{"type": "image", "image": Image.open("page.jpg")},
{"type": "text", "text": prompt},
]}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt", enable_thinking=False,
).to(model.device)
out = model.generate(
**inputs, max_new_tokens=8192, do_sample=False,
repetition_penalty=1.05, no_repeat_ngram_size=30,
)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```
Requirements: `transformers >= 5.10` or `vllm >= 0.24`, a GPU with 24 GB or
more for bf16. PDFs must be rasterized to page images first (around 150 to
200 DPI works well).
## Tips
- **Rotated photos.** The model expects upright pages. If your inputs may be
rotated, run an orientation check first. `orientation/` contains the small
CPU classifier (PP-LCNet_x1_0_doc_ori, 6.6 MB) and the script we use: rotate
only when the predicted angle has confidence 0.85 or higher.
- **Runaway output.** On rare pages with dense repeated structure the model can
keep generating until the token limit. If the output hits `max_tokens`,
regenerate with a slightly higher temperature (0.1 to 0.8) and keep the
shorter result.
- **Throughput.** One vLLM server per GPU with `--max-num-seqs 64` and 64
concurrent requests processes about 1.3 pages per second per GPU on
80 GB-class hardware.
## Model details
| | |
|---|---|
| Architecture | `Qwen3_5ForConditionalGeneration` |
| Parameters | 4.54 B |
| Weights | bf16 safetensors, 2 shards, 9.08 GB |
| Text decoder | 32 layers, hidden size 2560, vocabulary 248,320, context 262,144 |
| Vision encoder | 24 layers, patch 16, spatial merge 2 |
| Languages | de, en, es, fr, id, it, nl, pt, vi, ar, hi, ja, ko, ru, th, zh (Simplified and Traditional) |
| License | AGPL-3.0 (see `LICENSE`, `NOTICE`) |
## Benchmark
Official [MDPBench](https://github.com/Yuliang-Liu/MultimodalOCR/tree/main/MDPBench)
re-evaluation by the benchmark team (public set 2,720 pages, 17 languages,
digital-born and photographed):
| | Score |
|---|---|
| Public set, overall | 82.3 |
| Digital-born / Photographed | 87.3 / 80.7 |
| Private set, overall | 85.7 |
| Korean / Thai | **92.6 / 83.3** (highest on the leaderboard, September 2026) |
Per-language scores and the evaluation conditions are in
`eval_results/mdpbench.md`.
## Limitations
- Arabic, Japanese and French are the weakest of the supported languages.
- Tables come out as HTML, not Markdown pipe tables.
- One page per request; no multi-page or PDF input.
- Handwriting, very low-resolution scans and non-document photographs were not
part of training or evaluation targets.
## License
The weights and the files in this repository are released under the GNU Affero
General Public License v3.0 (`LICENSE`). You may use, modify and redistribute
them, including in networked services, provided the complete corresponding
source of your service is made available under the same license. For use under
different terms, including closed commercial deployment, contact Sionic AI.
PepperOCR-VL is a fine-tune of Qwen3.5-4B (Apache-2.0, Alibaba Cloud); the
optional orientation classifier is PP-LCNet_x1_0_doc_ori from PaddleOCR
(Apache-2.0). Their notices are preserved in `NOTICE`.
## Citation
```bibtex
@misc{pepperocr-vl-2026,
title = {PepperOCR-VL: Multilingual End-to-End Document Parsing},
author = {Sionic AI},
year = {2026},
howpublished = {\url{https://huggingface.co/sionic-ai/PepperOCR-VL}}
}
```