TeleOCR โ€” MLX INT8

MLX INT8 quantization of XingChen-AGI/TeleOCR, an OCR / document-parsing vision-language model. TeleOCR's config.json declares qwen2_5_vl, and its vision tower and multimodal RoPE are stock Qwen2.5-VL โ€” but its text decoder is not stock Qwen2.5-VL attention. The upstream modeling_naviocr.py (custom code shipped in the repo) patches every attention layer with per-head QK-RMSNorm (Qwen3-style q_norm/k_norm, applied before RoPE) and drops the qkv bias terms that vanilla Qwen2.5-VL uses โ€” verified directly against modeling_naviocr.py's Qwen2_5_VLAttention.forward (lines ~656-686) and the checkpoint's actual tensor shapes/keys, not just the config label.

This is why mlx-vlm's stock qwen2_5_vl module cannot load this checkpoint as-is (it has no q_norm/k_norm parameters and expects head_dim = hidden_size / num_heads = 64, but this model uses head_dim = 128). This repo ships a small Python patch, patch_qwen25vl_qknorm.py, that monkeypatches mlx-vlm's qwen2_5_vl.language.Attention (and TextConfig, to plumb the explicit head_dim) with a QK-norm-aware version at import time โ€” see Usage below. Runs on Apple Silicon via mlx-vlm once patched. Stays image-text-to-text โ€” the vision tower is kept in bf16; only the text backbone is quantized.

Precision INT8 (affine/group quant, group size 64)
Bits per weight 12.45 (text) / bf16 (vision)
On-disk size 1.8 GB (1 shard)
Quantized text backbone only (28 layers, hidden 1024, QK-normed attention, head_dim=128)
Kept in bf16 vision tower (stock Qwen2.5-VL ViT, depth 32, hidden 1280)
Embeddings tied (lm_head == embed_tokens); duplicate lm_head.weight dropped from the checkpoint before quantizing

Quantizations

Variant Size
TeleOCR-MXFP4 1.6 GB smaller / fastest
TeleOCR-MXFP8 1.8 GB higher fidelity
TeleOCR-INT8 1.8 GB โ† this repo

Verification

Tested with deterministic greedy decoding via native mlx-vlm, comparing against the unquantized bf16 reference (loaded through the same QK-norm patch):

  • Architecture sanity (text-only prompts): both the bf16 reference and all three quantized builds produce the same degenerate repeating output on generic chit-chat prompts ("capital of France", "25+17") โ€” TeleOCR is a narrow OCR/document specialist, not a general chat assistant, and this behavior is identical across precisions, confirming the QK-norm patch and quantization did not introduce new corruption (a broken patch would produce different garbage between bf16 and quantized, not matching garbage).
  • Real OCR task 1 (large, legible synthetic text, "Hello World 12345"): bf16, MXFP4, and MXFP8 all read it back exactly correct.
  • Real OCR task 2 (synthetic invoice image, Arial 28pt, "Invoice #A-9921" / "Total Due: $482.50"): bf16 reads both lines correctly (plus one spurious hallucinated prefix line); MXFP4 and MXFP8 read both lines correctly with no hallucination (cleaner than the bf16 reference on this sample); INT8 matches the bf16 reference exactly, hallucinated prefix included.
  • A low-quality/aliased-font synthetic image (PIL bitmap default font) caused the same unrelated hallucinated output across bf16 and all three quantized builds โ€” confirms the failure mode is image-legibility-driven, not quantization-driven (behavior is identical across precisions).

Usage (mlx-vlm)

pip install -U mlx-vlm

Download patch_qwen25vl_qknorm.py from this repo alongside the model, then:

import patch_qwen25vl_qknorm  # noqa: F401 -- must import before mlx_vlm.load()
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("sahilchachra/TeleOCR-INT8")
config = model.config

prompt = apply_chat_template(processor, config, "Read the text in this image.", num_images=1)
print(generate(model, processor, prompt, image=["invoice.png"], max_tokens=200, verbose=True))

Run in LM Studio

Not currently supported. LM Studio bundles its own mlx-vlm inside its MLX backend extension, and that bundled copy is the stock (un-patched) qwen2_5_vl module โ€” it does not know about this model's QK-norm attention or head_dim=128, and fails to load with ValueError: Received 56 parameters not in model: ...q_norm.weight, ...k_norm.weight (verified via lms load). Since the fix lives in a runtime monkeypatch of mlx-vlm's Python module (not something a GGUF/MLX file format flag can express), LM Studio cannot pick it up without XingChen-AGI's QK-norm variant being upstreamed into mlx-vlm itself. Use the native mlx-vlm path above instead.

Notes & limitations

  • The custom attention (QK-norm, no qkv bias, head_dim=128 despite hidden_size=1024/num_attention_heads=16) is not documented in TeleOCR's config.json model_type (qwen2_5_vl) or README โ€” it was found by diffing the checkpoint's tensor keys/shapes against mlx-vlm's stock qwen2_5_vl module and reading the shipped modeling_naviocr.py directly. If upstream TeleOCR changes its custom modeling code, this patch may need updating.
  • TeleOCR is trained as a document-OCR/parsing specialist; it is not tuned as a general chit-chat assistant, and generic non-document prompts can produce repetitive/degenerate output โ€” this is a base-model characteristic reproduced faithfully across all precisions here, not a quantization defect.
  • Inherits all capabilities and limitations of the base model and its Apache-2.0 license.
  • Quantized by @sahilchachra with MLX.
Downloads last month
38
Safetensors
Model size
1B params
Tensor type
U32
ยท
BF16
ยท
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sahilchachra/TeleOCR-INT8

Quantized
(4)
this model