jina-ocr-v1-8bit

This model was converted to MLX format from jinaai/jina-ocr-v1. Refer to the original model card for more details on the model.

Supported stacks

Stack Version Setup
mlx-vlm 0.7.4 Apply the fixes file before loading the model
omlx 0.7.0 Prompt suffix, --no-cache, sequential requests

Usage

Render each page 1400 px wide. A portrait page takes 926 prompt tokens whatever the render size, and a smaller render is upscaled by the processor.

mlx-vlm

Apply mlx_vlm_deepseekocr_fixes.py from this repository before loading the model and build the prompt with the chat template. model_path is the folder holding this repository's files:

import sys
sys.path.insert(0, model_path)
import mlx_vlm_deepseekocr_fixes
mlx_vlm_deepseekocr_fixes.apply_mlx_vlm_deepseekocr_fixes()

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load(model_path)
prompt = apply_chat_template(processor, model.config, "Transcribe the provided document image into a clean Markdown format, preserving the natural reading order.", num_images=1)
output = generate(model, processor, prompt, image=["page.png"], max_tokens=16000, temperature=0.0)
print(output.text)

omlx

The server does not load the file with the fixes. Apply these settings instead:

  • End the prompt with \n<|Assistant|>:\n. This supplies the missing assistant tag; the vision corrections are not available.

  • For reproducible output start the server with --no-cache and send requests sequentially. The flag disables the prefix cache for every model the server runs, so use a separate instance for OCR if other models should keep it.

Known issues

Port issues

mlx-vlm loads this model through its DeepSeek-OCR port and omlx uses that port as well. It departs from the reference implementation in two areas and some pages come back as invented text or endless repetition as a result.

  • Vision: the position embeddings of image tiles are resized differently from the reference and CLIP's MLP uses GELU where the reference uses QuickGELU.

  • Prompt: the model was trained on <|User|>:\n<image>\n{prompt}\n<|Assistant|>:\n. The port sends <image>\n{prompt}, prepends a BOS token and drops the last prompt token.

mlx_vlm_deepseekocr_fixes.py corrects all of these in memory when it is applied. Each correction checks the mlx-vlm source first and skips the patch once the issue is fixed upstream.

Prefix cache in omlx

omlx keeps prompt states in a prefix cache and restores them when a request repeats. A restored request can return different text than the first one at temperature 0 and each page adds about 150 MB to the cache.

Concurrent requests

Under concurrent requests the output is not reproducible at temperature 0. A long greedy decode is sensitive to small numerical differences between batched and single-request execution and a single differing token propagates through the rest of the sequence. On most pages nothing changes or the difference is confined to a few characters but on pages with large tables or charts it can be larger.

Measured performance

MacBook Pro, Apple M5 Pro, 48 GB, macOS 26.6.2. 92 pages, two chosen at random from each of 46 research papers, rendered 1400 px wide, greedy, through an omlx 0.7.0.dev4 build that applies the fixes in the server and does not cache OCR requests. Rendering time is not included.

Concurrent requests Seconds per page Output tokens per second Latency per page (median)
1 5.0 259 4.1 s
2 3.9 333 6.6 s
4 3.2 400 10.6 s
8 3.0 428 20.8 s

Each page takes 927 prompt tokens and a median of 1,100 output tokens. All 92 pages finished within the 16,000 token limit at every level of concurrency.

License

Quantized from jinaai/jina-ocr-v1, created by Jina AI (Alejandro Barón García, Feng Wang, Emilia Garcia Casademont, Han Xiao).

Modified from the original:

  • The weights are quantized to 8 bits.

  • The FastMTP head is removed, both its weights and its mtp_* entries in config.json.

  • config.json declares model_type as deepseekocr instead of deepseek_vl_v2 and has no auto_map entries.

  • processor_config.json is rewritten in the format of mlx-vlm's DeepSeek-OCR processor.

  • tokenizer_config.json has no auto_map entry and names TokenizersBackend as its tokenizer_class instead of LlamaTokenizerFast.

  • The PyTorch files are removed: configuration_deepseek_v2.py, deepencoder.py, deepseek_ocr_mtp.py, example.py, modeling_deepseekocr.py, modeling_deepseekv2.py and processing_deepseek_ocr.py.

  • In chat_template.jinja the double-escaped newlines are corrected, so the template renders line breaks instead of a literal backslash and n.

Licensed under CC BY-NC 4.0, the same license as the original. Commercial use: contact Jina AI.

mlx_vlm_deepseekocr_fixes.py is not part of the original model and is licensed separately under Apache 2.0.

Downloads last month
135
Safetensors
Model size
3B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stefanschmidt/jina-ocr-v1-8bit

Quantized
(3)
this model