manga-ocr-base-mlx

English | 日本語

Model Summary

This is an unofficial MLX conversion of kha-white/manga-ocr-base (a Japanese OCR model specialized for manga speech-bubble text). All credit for the original model goes to its author.

This cannot be loaded with mlx-lm / mlx-embeddings / mlx-vlm

This model is a VisionEncoderDecoderModel (ViT encoder + BERT-style cross-attention decoder), which isn't in mlx-vlm's list of supported architectures either. So it was reimplemented from scratch for MLX and requires the bundled manga_ocr_mlx.py.

Usage

pip install mlx transformers fugashi unidic-lite pillow
from huggingface_hub import snapshot_download
import sys

path = snapshot_download("masahiroid/manga-ocr-base-mlx")
sys.path.insert(0, path)

import mlx.core as mx
from manga_ocr_mlx import MangaOcrMLX
from transformers import ViTImageProcessor, AutoTokenizer
from PIL import Image

model = MangaOcrMLX()
model.load_weights(f"{path}/model.safetensors")
mx.eval(model.parameters())

processor = ViTImageProcessor.from_pretrained("kha-white/manga-ocr-base")
tokenizer = AutoTokenizer.from_pretrained("kha-white/manga-ocr-base")

img = Image.open("page.png").convert("RGB")
pixel_values = mx.array(processor(img, return_tensors="np").pixel_values)

encoder_hidden_states = model.encode(pixel_values)

decoder_start_token_id = 2
eos_token_id = 3
generated = [decoder_start_token_id]
for _ in range(30):
    logits = model.decode(mx.array([generated]), encoder_hidden_states)
    next_id = int(mx.argmax(logits[0, -1]).item())
    generated.append(next_id)
    if next_id == eos_token_id:
        break

print(tokenizer.decode(generated, skip_special_tokens=True))

No KV cache is implemented (since the decoder is only 2 layers deep, recomputing the whole sequence each step is still fast enough).

Accuracy

Verified on 2 generated test images ("瑠璃色の空", "ありがとうございます") that output exactly matches PyTorch's model.generate() output token sequence.

Notes

  • The original model (kha-white/manga-ocr-base) is distributed only as pytorch_model.bin (pickle format); this converted version consists only of safetensors.
  • This is a community conversion, not an official release from the original author.

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

kha-white/manga-ocr-base(漫画のフキダシ文字認識に特化した日本語OCRモデル)のMLX版です。元モデルの著作権はその作者に帰属します。

これはmlx-lm/mlx-embeddings/mlx-vlmでは読み込めません

このモデルはVisionEncoderDecoderModel(ViTエンコーダー + BERT系クロスアテンションデコーダー)という 構成で、mlx-vlmの対応アーキテクチャ一覧にも含まれていません。そのため、MLXでの実装をゼロから 書き起こして変換しています。同梱のmanga_ocr_mlx.pyが必要です。

使い方

pip install mlx transformers fugashi unidic-lite pillow
from huggingface_hub import snapshot_download
import sys

path = snapshot_download("masahiroid/manga-ocr-base-mlx")
sys.path.insert(0, path)

import mlx.core as mx
from manga_ocr_mlx import MangaOcrMLX
from transformers import ViTImageProcessor, AutoTokenizer
from PIL import Image

model = MangaOcrMLX()
model.load_weights(f"{path}/model.safetensors")
mx.eval(model.parameters())

processor = ViTImageProcessor.from_pretrained("kha-white/manga-ocr-base")
tokenizer = AutoTokenizer.from_pretrained("kha-white/manga-ocr-base")

img = Image.open("page.png").convert("RGB")
pixel_values = mx.array(processor(img, return_tensors="np").pixel_values)

encoder_hidden_states = model.encode(pixel_values)

decoder_start_token_id = 2
eos_token_id = 3
generated = [decoder_start_token_id]
for _ in range(30):
    logits = model.decode(mx.array([generated]), encoder_hidden_states)
    next_id = int(mx.argmax(logits[0, -1]).item())
    generated.append(next_id)
    if next_id == eos_token_id:
        break

print(tokenizer.decode(generated, skip_special_tokens=True))

KVキャッシュは実装していません(デコーダーが2層と浅いため、毎回全系列を再計算しても十分高速です)。

精度検証

生成したテスト画像2種(「瑠璃色の空」「ありがとうございます」)で、PyTorchのmodel.generate()の 出力トークン列と完全に一致することを確認しています。

備考

  • 元モデル(kha-white/manga-ocr-base)はpytorch_model.bin(pickle形式)のみで配布されていますが、 本変換版はsafetensorsのみで構成されています。
  • 本変換は非公式のコミュニティ版です。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.1B params
Tensor type
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/manga-ocr-base-mlx

Finetuned
(3)
this model