manga-ocr-base-coreml

English | 日本語

Model Summary

This is an unofficial Core ML conversion of kha-white/manga-ocr-base (a Japanese OCR model specialized for manga speech-bubble text), for running directly on iOS/macOS. All credit for the original model goes to its author.

Architecture

A ViT encoder (image -> features) + a 2-layer BERT decoder (features -> text, autoregressive, character-level). Since the decoder is only 2 layers deep, a simple design without a KV cache still runs at practical speed.

File Content Size
manga-ocr-base_encoder_fp16.mlpackage Image encoder (fixed input: 224x224 RGB) ~164MB
manga-ocr-base_decoder_seq32_fp16.mlpackage Text decoder (sequence length 32) ~46MB
manga-ocr-base_decoder_seq64_fp16.mlpackage Text decoder (sequence length 64, for longer lines) ~46MB

Usage

1. Image preprocessing

Resize and normalize the image to 224x224. Use transformers' ViTImageProcessor (kha-white/manga-ocr-base), or an equivalent implementation.

2. Run the encoder

encoder_hidden_states = encoder.predict({"pixel_values": pixel_values})["encoder_hidden_states"]

3. Run the decoder (greedy decoding, no KV cache)

decoder_start_token_id = 2
eos_token_id = 3
SEQ_LEN = 32  # match the decoder file you're using

generated = [decoder_start_token_id]
for step in range(SEQ_LEN - 1):
    padded = generated + [0] * (SEQ_LEN - len(generated))  # pad with pad_token_id=0
    logits = decoder.predict({
        "decoder_input_ids": np.array([padded], dtype=np.int32),
        "encoder_hidden_states": encoder_hidden_states.astype(np.float16),
    })["logits"]
    next_id = int(np.argmax(logits[0, len(generated) - 1]))
    generated.append(next_id)
    if next_id == eos_token_id:
        break

text = tokenizer.decode(generated, skip_special_tokens=True)

Uses a character-level tokenizer (BertJapaneseTokenizer, based on cl-tohoku/bert-base-japanese-char-v2).

Accuracy

Verified on 2 generated test images ("瑠璃色の空", "ありがとうございます") that the Core ML output token sequence (via the procedure above) exactly matches PyTorch's model.generate() output token sequence.

Notes

  • The original model (kha-white/manga-ocr-base) is distributed only as pytorch_model.bin (pickle format); this converted version consists only of safetensors.
  • This is a community conversion, not an official release from the original author.

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

kha-white/manga-ocr-base(漫画のフキダシ文字認識に特化した日本語OCRモデル)を、iOS/macOS (Core ML) で直接動かせるように変換したものです。元モデルの著作権はその作者に帰属します。

構成

ViTエンコーダー(画像→特徴量)+ 2層BERTデコーダー(特徴量→テキスト、自己回帰・文字単位)という構成です。 デコーダーが2層と浅いため、KVキャッシュなしのシンプルな設計で実用的な速度が出ます。

ファイル 内容 サイズ
manga-ocr-base_encoder_fp16.mlpackage 画像エンコーダー(固定入力: 224×224 RGB) 約164MB
manga-ocr-base_decoder_seq32_fp16.mlpackage テキストデコーダー(系列長32) 約46MB
manga-ocr-base_decoder_seq64_fp16.mlpackage テキストデコーダー(系列長64、長めのセリフ向け) 約46MB

使い方

1. 画像の前処理

画像を224×224にリサイズ・正規化します。transformersのViTImageProcessor(kha-white/manga-ocr-base)、 または同等の前処理を実装してください。

2. エンコーダーの実行

encoder_hidden_states = encoder.predict({"pixel_values": pixel_values})["encoder_hidden_states"]

3. デコーダーの実行(グリーディデコード、KVキャッシュなし)

decoder_start_token_id = 2
eos_token_id = 3
SEQ_LEN = 32  # 使用するデコーダーのファイルに合わせる

generated = [decoder_start_token_id]
for step in range(SEQ_LEN - 1):
    padded = generated + [0] * (SEQ_LEN - len(generated))  # pad_token_id=0 で埋める
    logits = decoder.predict({
        "decoder_input_ids": np.array([padded], dtype=np.int32),
        "encoder_hidden_states": encoder_hidden_states.astype(np.float16),
    })["logits"]
    next_id = int(np.argmax(logits[0, len(generated) - 1]))
    generated.append(next_id)
    if next_id == eos_token_id:
        break

text = tokenizer.decode(generated, skip_special_tokens=True)

文字単位トークナイザ(BertJapaneseTokenizer, cl-tohoku/bert-base-japanese-char-v2ベース)を使用します。

精度検証

生成したテスト画像2種(「瑠璃色の空」「ありがとうございます」)で、PyTorchのmodel.generate()の 出力トークン列と、上記手順によるCore ML版の出力トークン列が完全に一致することを確認しています。

備考

  • 元モデル(kha-white/manga-ocr-base)はpytorch_model.bin(pickle形式)のみで配布されていますが、 本変換版はsafetensorsのみで構成されています。
  • 本変換は非公式のコミュニティ版です。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/manga-ocr-base-coreml

Quantized
(8)
this model