nemu on-device manga OCR (Core ML)

These are the Core ML models that nemu's iOS app bundles for fully on-device Japanese manga OCR. They power the reader's Japanese-learning plugin. The pipeline has three stages:

  1. Detect. ogkalu-ctbd-v4s-fp16.mlpackage is a text and speech-bubble detector: RT-DETRv4-S from ogkalu/comic-text-and-bubble-detector, with a 640×640 input. It finds the text blocks on a page. It was built from detector-v4-s_int8.onnx at revision 16e8a622: dequantised, then converted to an fp16 ML Program.
  2. Order. Blocks are put into manga reading order in the app. This step is code, not a model.
  3. Recognise. mangaocr-encoder-int8.mlpackage and mangaocr-decoder-int8.mlpackage read each block. They are kha-white/manga-ocr-base split into a ViT encoder and a 2-layer BERT decoder, which the host runs step by step with greedy decoding. Weights are int8 linear, per channel; activations stay fp16. vocab.txt is the upstream tokenizer vocabulary.
file size notes
ogkalu-ctbd-v4s-fp16.mlpackage ~20 MB raw per-query outputs: logits [1,300,3] (bubble, text_bubble, text_free) and boxes [1,300,4] (normalised cx, cy, w, h). Post-processing runs in Swift.
mangaocr-encoder-int8.mlpackage ~82 MB input: a 224×224 RGB crop, preprocessed with an antialiased bilinear resize to match PIL.
mangaocr-decoder-int8.mlpackage ~24 MB flexible input_ids length 1–300.
vocab.txt 24 KB

Every package targets iOS 17 or later and runs on the Neural Engine, GPU or CPU. In the iOS Simulator, use .cpuOnly.

Quality

The benchmark set is 16 raw manga pages with 148 hand-transcribed dialogue and narration blocks. The fully on-device pipeline, measured in the nemu app, scored:

metric score
Normalised CER 1.1% (Apple Vision baseline: 18.7%)
Block recall 98%
Exact blocks 95%
Pages in correct reading order 15/16

For recognition, the Swift implementation's output matches the Python Core ML reference on 166/166 crops. The int8 model matches upstream PyTorch greedy decoding on 164/166 crops.

Reproducing

The scripts in scripts/ rebuild the packages, using coremltools 9, torch, transformers and onnx2torch:

export NEMU_OCR_WORKDIR=$PWD/work   # expects work/models/<upstream weights>, writes work/coreml/
python scripts/convert_mangaocr.py fp16          # manga-ocr → encoder/decoder fp16
python scripts/quantize_int8.py                  # fp16 → int8 weights
python scripts/dequantize_int8_onnx.py detector-v4-s_int8.onnx detector-v4-s.fp32.onnx
python scripts/convert_v4s_coreml.py detector-v4-s.fp32.onnx work/coreml/ogkalu-ctbd-v4s

The nemu build pins this repository by commit and checks every file against a sha256 hash (apps/mobile/modules/nemu-japanese-learning/scripts/fetch-ocr-models.ts). It then compiles the packages with xcrun coremlcompiler and bundles them into the app.

Licence and attribution

Both upstream models are released under Apache-2.0, and so are these converted derivatives. See NOTICE for full attribution.

  • manga-ocr-base: © Maciej Budyś (kha-white). It was trained on Manga109-s (Aizawa-Yamasaki-Matsui Laboratory, The University of Tokyo) and synthetic data. No Manga109 images are included here.
  • comic-text-and-bubble-detector: © ogkalu. The model card describes ~11k manga, webtoon, manhua and comic training images; it does not publish a full list of image sources.

Changes made by nemu: conversion to Core ML, splitting the encoder and decoder, int8 weight quantisation of the recogniser, and dequantising and cutting the detector's ONNX graph before its post-processor.

Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nemu-pm/nemu-ocr-coreml

Quantized
(8)
this model