Instructions to use stefanschmidt/jina-ocr-v1-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use stefanschmidt/jina-ocr-v1-8bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("stefanschmidt/jina-ocr-v1-8bit") config = load_config("stefanschmidt/jina-ocr-v1-8bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
jina-ocr-v1-8bit
This model was converted to MLX format from jinaai/jina-ocr-v1. Refer to the original model card for more details on the model.
Supported stacks
| Stack | Version | Setup |
|---|---|---|
| mlx-vlm | 0.7.4 | Apply the fixes file before loading the model |
| omlx | 0.7.0 | Prompt suffix, --no-cache, sequential requests |
Usage
Render each page 1400 px wide. A portrait page takes 926 prompt tokens whatever the render size, and a smaller render is upscaled by the processor.
mlx-vlm
Apply mlx_vlm_deepseekocr_fixes.py from this repository before loading the model and build the prompt with the chat template. model_path is the folder holding this repository's files:
import sys
sys.path.insert(0, model_path)
import mlx_vlm_deepseekocr_fixes
mlx_vlm_deepseekocr_fixes.apply_mlx_vlm_deepseekocr_fixes()
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, processor = load(model_path)
prompt = apply_chat_template(processor, model.config, "Transcribe the provided document image into a clean Markdown format, preserving the natural reading order.", num_images=1)
output = generate(model, processor, prompt, image=["page.png"], max_tokens=16000, temperature=0.0)
print(output.text)
omlx
The server does not load the file with the fixes. Apply these settings instead:
End the prompt with
\n<|Assistant|>:\n. This supplies the missing assistant tag; the vision corrections are not available.For reproducible output start the server with
--no-cacheand send requests sequentially. The flag disables the prefix cache for every model the server runs, so use a separate instance for OCR if other models should keep it.
Known issues
Port issues
mlx-vlm loads this model through its DeepSeek-OCR port and omlx uses that port as well. It departs from the reference implementation in two areas and some pages come back as invented text or endless repetition as a result.
Vision: the position embeddings of image tiles are resized differently from the reference and CLIP's MLP uses GELU where the reference uses QuickGELU.
Prompt: the model was trained on
<|User|>:\n<image>\n{prompt}\n<|Assistant|>:\n. The port sends<image>\n{prompt}, prepends a BOS token and drops the last prompt token.
mlx_vlm_deepseekocr_fixes.py corrects all of these in memory when it is applied. Each correction checks the mlx-vlm source first and skips the patch once the issue is fixed upstream.
Prefix cache in omlx
omlx keeps prompt states in a prefix cache and restores them when a request repeats. A restored request can return different text than the first one at temperature 0 and each page adds about 150 MB to the cache.
Concurrent requests
Under concurrent requests the output is not reproducible at temperature 0. A long greedy decode is sensitive to small numerical differences between batched and single-request execution and a single differing token propagates through the rest of the sequence. On most pages nothing changes or the difference is confined to a few characters but on pages with large tables or charts it can be larger.
Measured performance
MacBook Pro, Apple M5 Pro, 48 GB, macOS 26.6.2. 92 pages, two chosen at random from each of 46 research papers, rendered 1400 px wide, greedy, through an omlx 0.7.0.dev4 build that applies the fixes in the server and does not cache OCR requests. Rendering time is not included.
| Concurrent requests | Seconds per page | Output tokens per second | Latency per page (median) |
|---|---|---|---|
| 1 | 5.0 | 259 | 4.1 s |
| 2 | 3.9 | 333 | 6.6 s |
| 4 | 3.2 | 400 | 10.6 s |
| 8 | 3.0 | 428 | 20.8 s |
Each page takes 927 prompt tokens and a median of 1,100 output tokens. All 92 pages finished within the 16,000 token limit at every level of concurrency.
License
Quantized from jinaai/jina-ocr-v1, created by Jina AI (Alejandro Barón García, Feng Wang, Emilia Garcia Casademont, Han Xiao).
Modified from the original:
The weights are quantized to 8 bits.
The FastMTP head is removed, both its weights and its
mtp_*entries inconfig.json.config.jsondeclaresmodel_typeasdeepseekocrinstead ofdeepseek_vl_v2and has noauto_mapentries.processor_config.jsonis rewritten in the format of mlx-vlm's DeepSeek-OCR processor.tokenizer_config.jsonhas noauto_mapentry and namesTokenizersBackendas itstokenizer_classinstead ofLlamaTokenizerFast.The PyTorch files are removed:
configuration_deepseek_v2.py,deepencoder.py,deepseek_ocr_mtp.py,example.py,modeling_deepseekocr.py,modeling_deepseekv2.pyandprocessing_deepseek_ocr.py.In
chat_template.jinjathe double-escaped newlines are corrected, so the template renders line breaks instead of a literal backslash andn.
Licensed under CC BY-NC 4.0, the same license as the original. Commercial use: contact Jina AI.
mlx_vlm_deepseekocr_fixes.py is not part of the original model and is licensed separately under Apache 2.0.
- Downloads last month
- 135
8-bit
Model tree for stefanschmidt/jina-ocr-v1-8bit
Base model
jinaai/jina-ocr-v1