Instructions to use masahiroid/manga-ocr-base-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use masahiroid/manga-ocr-base-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir manga-ocr-base-mlx masahiroid/manga-ocr-base-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
manga-ocr-base-mlx
Model Summary
This is an unofficial MLX conversion of kha-white/manga-ocr-base (a Japanese OCR model specialized for manga speech-bubble text). All credit for the original model goes to its author.
This cannot be loaded with mlx-lm / mlx-embeddings / mlx-vlm
This model is a VisionEncoderDecoderModel (ViT encoder + BERT-style
cross-attention decoder), which isn't in mlx-vlm's list of supported
architectures either. So it was reimplemented from scratch for MLX and
requires the bundled manga_ocr_mlx.py.
Usage
pip install mlx transformers fugashi unidic-lite pillow
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("masahiroid/manga-ocr-base-mlx")
sys.path.insert(0, path)
import mlx.core as mx
from manga_ocr_mlx import MangaOcrMLX
from transformers import ViTImageProcessor, AutoTokenizer
from PIL import Image
model = MangaOcrMLX()
model.load_weights(f"{path}/model.safetensors")
mx.eval(model.parameters())
processor = ViTImageProcessor.from_pretrained("kha-white/manga-ocr-base")
tokenizer = AutoTokenizer.from_pretrained("kha-white/manga-ocr-base")
img = Image.open("page.png").convert("RGB")
pixel_values = mx.array(processor(img, return_tensors="np").pixel_values)
encoder_hidden_states = model.encode(pixel_values)
decoder_start_token_id = 2
eos_token_id = 3
generated = [decoder_start_token_id]
for _ in range(30):
logits = model.decode(mx.array([generated]), encoder_hidden_states)
next_id = int(mx.argmax(logits[0, -1]).item())
generated.append(next_id)
if next_id == eos_token_id:
break
print(tokenizer.decode(generated, skip_special_tokens=True))
No KV cache is implemented (since the decoder is only 2 layers deep, recomputing the whole sequence each step is still fast enough).
Accuracy
Verified on 2 generated test images ("瑠璃色の空", "ありがとうございます") that
output exactly matches PyTorch's model.generate() output token sequence.
Notes
- The original model (
kha-white/manga-ocr-base) is distributed only aspytorch_model.bin(pickle format); this converted version consists only ofsafetensors. - This is a community conversion, not an official release from the original author.
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
kha-white/manga-ocr-base(漫画のフキダシ文字認識に特化した日本語OCRモデル)のMLX版です。元モデルの著作権はその作者に帰属します。
これはmlx-lm/mlx-embeddings/mlx-vlmでは読み込めません
このモデルはVisionEncoderDecoderModel(ViTエンコーダー + BERT系クロスアテンションデコーダー)という
構成で、mlx-vlmの対応アーキテクチャ一覧にも含まれていません。そのため、MLXでの実装をゼロから
書き起こして変換しています。同梱のmanga_ocr_mlx.pyが必要です。
使い方
pip install mlx transformers fugashi unidic-lite pillow
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("masahiroid/manga-ocr-base-mlx")
sys.path.insert(0, path)
import mlx.core as mx
from manga_ocr_mlx import MangaOcrMLX
from transformers import ViTImageProcessor, AutoTokenizer
from PIL import Image
model = MangaOcrMLX()
model.load_weights(f"{path}/model.safetensors")
mx.eval(model.parameters())
processor = ViTImageProcessor.from_pretrained("kha-white/manga-ocr-base")
tokenizer = AutoTokenizer.from_pretrained("kha-white/manga-ocr-base")
img = Image.open("page.png").convert("RGB")
pixel_values = mx.array(processor(img, return_tensors="np").pixel_values)
encoder_hidden_states = model.encode(pixel_values)
decoder_start_token_id = 2
eos_token_id = 3
generated = [decoder_start_token_id]
for _ in range(30):
logits = model.decode(mx.array([generated]), encoder_hidden_states)
next_id = int(mx.argmax(logits[0, -1]).item())
generated.append(next_id)
if next_id == eos_token_id:
break
print(tokenizer.decode(generated, skip_special_tokens=True))
KVキャッシュは実装していません(デコーダーが2層と浅いため、毎回全系列を再計算しても十分高速です)。
精度検証
生成したテスト画像2種(「瑠璃色の空」「ありがとうございます」)で、PyTorchのmodel.generate()の
出力トークン列と完全に一致することを確認しています。
備考
- 元モデル(
kha-white/manga-ocr-base)はpytorch_model.bin(pickle形式)のみで配布されていますが、 本変換版はsafetensorsのみで構成されています。 - 本変換は非公式のコミュニティ版です。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
Quantized
Model tree for masahiroid/manga-ocr-base-mlx
Base model
kha-white/manga-ocr-base