TeleOCR โ€” GGUF for teleocr-rs

GGUF conversions of XingChen-AGI/TeleOCR (China Telecom, ~1.2B, document parsing: layout, text, LaTeX formulas, tables, code) for teleocr-rs, a Rust/Candle port.

Not for llama.cpp. TeleOCR's decoder is Qwen3-style (head_dim 128 โ‰  hidden/heads, per-head QK-norm) behind a Qwen2.5-VL config, which llama.cpp does not load. These files are single-file GGUFs (weights + config + tokenizer) in the teleocr-rs layout.

file size weights use
teleocr-q8v.gguf 1.5 GB text and vision linears Q8_0 GPU (recommended)
teleocr-q8_0.gguf 2.0 GB text Q8_0, vision F16 CPU (Q8_0 vision is slow on CPU)

Accuracy

Greedy output is token-identical to the transformers reference on the five model-card samples (text, formula, table, code, 767-token layout), for both files, on CPU and CUDA. Both files also give identical Markdown on six pages of the TeleOCR paper.

Speed (RTX 4090)

~225 tokens/s per sequence; vision tower 0.38 s for a 1036ร—1036 page; full page parsing (layout + all blocks, batch 8) ~3.2 s per page. Peak VRAM ~6.4 GB at batch 8.

Usage

git clone https://github.com/rzafiamy/teleocr-rs && cd teleocr-rs
cargo build --release -p teleocr-cli --features cuda   # or without --features cuda for CPU
scripts/fetch-pdfium.sh target/release                  # PDF input

target/release/teleocr parse -m teleocr-q8v.gguf paper.pdf --pages 1-3       # Markdown
target/release/teleocr run   -m teleocr-q8v.gguf table.png -t table           # one task
target/release/teleocr serve -m teleocr-q8v.gguf --port 8090                  # POST /v1/ocr

Also served by zallama (modality: ocr, POST /v1/ocr).

Reproduce

teleocr convert ./TeleOCR -o teleocr-q8v.gguf --vision-dtype q8_0
teleocr convert ./TeleOCR -o teleocr-q8_0.gguf            # vision F16

License

Apache-2.0, as the original model. Credits: Cai et al., TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents, arXiv:2608.12898.

Downloads last month
-
GGUF
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for rleo/TeleOCR-GGUF

Quantized
(5)
this model

Paper for rleo/TeleOCR-GGUF