TeleOCR โ GGUF for teleocr-rs
GGUF conversions of XingChen-AGI/TeleOCR (China Telecom, ~1.2B, document parsing: layout, text, LaTeX formulas, tables, code) for teleocr-rs, a Rust/Candle port.
Not for llama.cpp. TeleOCR's decoder is Qwen3-style (head_dim 128 โ hidden/heads, per-head QK-norm) behind a Qwen2.5-VL config, which llama.cpp does not load. These files are single-file GGUFs (weights + config + tokenizer) in the teleocr-rs layout.
| file | size | weights | use |
|---|---|---|---|
teleocr-q8v.gguf |
1.5 GB | text and vision linears Q8_0 | GPU (recommended) |
teleocr-q8_0.gguf |
2.0 GB | text Q8_0, vision F16 | CPU (Q8_0 vision is slow on CPU) |
Accuracy
Greedy output is token-identical to the transformers reference on the five model-card samples (text, formula, table, code, 767-token layout), for both files, on CPU and CUDA. Both files also give identical Markdown on six pages of the TeleOCR paper.
Speed (RTX 4090)
~225 tokens/s per sequence; vision tower 0.38 s for a 1036ร1036 page; full page parsing (layout + all blocks, batch 8) ~3.2 s per page. Peak VRAM ~6.4 GB at batch 8.
Usage
git clone https://github.com/rzafiamy/teleocr-rs && cd teleocr-rs
cargo build --release -p teleocr-cli --features cuda # or without --features cuda for CPU
scripts/fetch-pdfium.sh target/release # PDF input
target/release/teleocr parse -m teleocr-q8v.gguf paper.pdf --pages 1-3 # Markdown
target/release/teleocr run -m teleocr-q8v.gguf table.png -t table # one task
target/release/teleocr serve -m teleocr-q8v.gguf --port 8090 # POST /v1/ocr
Also served by zallama (modality: ocr, POST /v1/ocr).
Reproduce
teleocr convert ./TeleOCR -o teleocr-q8v.gguf --vision-dtype q8_0
teleocr convert ./TeleOCR -o teleocr-q8_0.gguf # vision F16
License
Apache-2.0, as the original model. Credits: Cai et al., TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents, arXiv:2608.12898.
- Downloads last month
- -
8-bit
Model tree for rleo/TeleOCR-GGUF
Base model
XingChen-AGI/TeleOCR