Unlimited-OCR, served through MAX
baidu/Unlimited-OCR served through
MAX as an OpenAI-compatible endpoint on Metal,
CUDA, ROCm or CPU. Code: github.com/kthr/unlimited-ocr-max.
- bf16 — baidu's
model.safetensors, unchanged (upstream revision07dea832e22aefee32ad281d4b80551282e1c168, sha2562bc48a7a110061ea58fff65d3169367eebe3aee371ca6968dc2219c1b2855fc6). - int8 —
model-int8.safetensors, this port's weight-only quantisation of the routed experts (symmetric, per-group G=128). config.jsonis baidu's minusauto_mapandmodel_type, so MAX loads it without remote code.
Prerequisites
Python 3.12 or 3.13, plus one of:
Apple silicon — M1–M5, macOS 15+, 24 GB recommended
# requires full Xcode (App Store); the Command Line Tools alone are not enough
sudo xcode-select -s /Applications/Xcode.app/Contents/Developer
# MAX compiles Metal kernels with `xcrun metallib`, shipped in the Metal Toolchain
xcodebuild -downloadComponent MetalToolchain
xcrun -f metallib # must print a path
NVIDIA — Linux (glibc 2.34+, e.g. Ubuntu 22.04+), Ampere or newer, driver 580+
Follow NVIDIA's package-manager installation; for example, on Ubuntu 24.04 x86_64:
# MAX loads these at runtime but does not ship them; without them the
# first request fails with: symbol not found: cublasCreate_v2
# 1. add NVIDIA's CUDA repository (ubuntu2404/x86_64 shown; use your distro/arch)
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
# 2. libcublas + libcublasLt (one package), CUDA 13
sudo apt-get install -y libcublas-13-0
# 3. check: MAX loads these exact files
ls -l /usr/local/cuda-13.0/lib64/libcublas.so.13 /usr/local/cuda-13.0/lib64/libcublasLt.so.13
Use these system packages, not the nvidia-*-cu13 pip wheels — the loader does
not find those. NVIDIA needs v0.3.1 or later: v0.3.0 aborts on the first
request with CUDA_ERROR_INVALID_VALUE.
AMD — Linux, driver 6.3.3+ (MI355X: ROCm 7+)
MAX also loads ROCm's rocBLAS, hipBLASLt and MIOpen from /opt/rocm/lib;
untested beyond compilation.
CPU
Nothing beyond Python. Supported, slow.
Serve
uv tool install unlimited-ocr-max # or: pip install unlimited-ocr-max
unlimited-ocr-max serve --devices gpu
This installs max[all]==26.6.0, numpy and pillow from PyPI (no extra
index). The first run downloads the weights (6.2 GiB bf16, 4.0 GiB int8) and
compiles the kernels. Endpoint: http://127.0.0.1:8010/v1/chat/completions, model
unlimited-ocr-max.
| flag | values | default | what it does |
|---|---|---|---|
--devices |
gpu | cpu |
required | gpu is Metal, CUDA or ROCm; cpu is slow |
--weights |
bf16 | int8 |
bf16 |
int8: quantised routed experts, faster decode, gpu only |
--model |
Hub repo or local dir | kthierbach/unlimited-ocr-max |
a local dir needs this repository's layout |
--revision |
tag | v0.3.2 |
the model-repo tag this package version was validated against |
--port |
integer | 8010 |
|
--ngram-size |
integer | 35 |
no-repeat n-gram guard; 0 disables it |
Measure it on your machine
unlimited-ocr-max profile --devices gpu # add --weights int8, or --devices cpu
Starts its own server, sends the 12 bundled pages, prints one row of the table
below and writes profile.json. On a GPU it refuses to run unless no other
process is using the GPU. --out DIR sets the output directory (default
./unlimited-ocr-max-profile-<UTC time>).
One page to Markdown:
curl -s http://127.0.0.1:8010/v1/chat/completions -H 'Content-Type: application/json' -d @- <<EOF
{"model": "unlimited-ocr-max", "temperature": 0, "max_tokens": 1766,
"messages": [{"role": "user", "content": [
{"type": "text", "text": "<|grounding|>Convert the document to markdown."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,$(base64 < page.png | tr -d '\n')"}}]}]}
EOF
Offline:
uvx --from huggingface_hub hf download kthierbach/unlimited-ocr-max --revision v0.3.2 --local-dir ocr-model
unlimited-ocr-max serve --devices gpu --model ocr-model
Where it has run
temperature 0, default guard. Text is compared against the fp32 PyTorch
reference (transformers 4.46.3, CPU); CER = edited characters / reference
characters over all pages.
| hardware | weights | status | decode | prefill | memory, peak / steady | text vs reference |
|---|---|---|---|---|---|---|
| Apple M4 24 GB | bf16 | 12 pages | 21.3 tok/s | 3.20 s | 17.8 / 11.1–11.2 GiB | 12/12 byte-identical |
| Apple M4 24 GB | int8 | 12 pages ² | 36.6 tok/s | 6.86 s | 20.7 / 10.9 GiB | 6/12; CER 0.0011, all edits bbox digits |
| NVIDIA A100 80 GB | bf16 | 12 pages | 93.6 tok/s | 3.56 s | device 10.6 / 10.6 GiB | 10/12; CER 0.022 ³ |
| NVIDIA A100 80 GB | int8 | 12 pages | 82.7 tok/s | 6.23 s | device 17.4 / 17.4 GiB | 6/12; CER 0.025 ³ |
| NVIDIA T4 (Turing, sm_75) | any | does not run ¹ | — | — | — | — |
| AMD gfx90a / gfx942 / gfx950 / gfx1100 | both | compiles, never served | — | — | — | — |
| CPU (Apple M4) | bf16 | 12 pages ² | 5.1 tok/s | 47.9 s | 21.1 / 21.0 GiB | 12/12 byte-identical |
Every measured row is one draw with unlimited-ocr-max profile on v0.3.2
(Apple M4: macOS 26.5.2; A100: Linux; max 26.6.0 on both). Memory is the server's physical
footprint on Apple silicon -- what footprint and vmmap report, which on
unified memory includes the Metal allocations and excludes clean page cache --
and its device memory on NVIDIA/AMD. Peak is reached in the first request
(weight upload and graph compile); steady is the server idle between pages.
¹ Upstream: MAX's ldmatrix PTX needs sm_80, and Turing has no bf16 tensor
cores (modular/modular#6653,
#6659). ² System swap grew
during these runs, so their timings are indicative; the text results are
exact. ³ On CUDA two of the twelve pages differ from the fp32 reference. The
bf16 output is byte-identical to v0.3.1's on the same A100, so this is CUDA's
arithmetic rather than a change in this release.
Not supported
gundam (tiled) mode; batch size > 1; multi-GPU.
References
Unlimited OCR Works (Yin et al., 2026), arXiv:2606.23050.
This port changes how the model is served, not the model (int8 excepted, as
quantified above); model-level behaviour, limitations and biases are those of
baidu/Unlimited-OCR.
@misc{yin2026unlimitedocrworks,
title={Unlimited OCR Works},
author={Youyang Yin and Huanhuan Liu and YY and Qunyi Xie and Chaorun Liu and Shiqi Yang and Shaohua Wang and Zhanlong Liu and Hao Zou and Jinyue Chen and Shu Wei and Jingjing Wu and Mingxin Huang and Zhen Wu and Guibin Wang and Tengyu Du and Lei Jia},
year={2026},
eprint={2606.23050},
archivePrefix={arXiv}
}
Authorship
This port, its serve configuration and this documentation were written with substantial AI assistance (Claude, via Claude Code); Konstantin Thierbach reviewed them and is accountable for their contents. Assisted-by: AI.
License
MIT — see LICENSE: this port's MIT notice and baidu's verbatim. The
bf16 weights, tokenizer files and config.json are baidu's, from
baidu/Unlimited-OCR (MIT, Copyright (c) 2026 Baidu), redistributed under that
license; the bf16 weights are unchanged, config.json has two keys removed, and
model-int8.safetensors is derived from those weights by this port's quantiser.
The profile corpus bundles renders of four pages of that paper
(arXiv:2606.23050); see unlimited_ocr_max/profile_data/NOTICE.md.
- Downloads last month
- 279