JMangaTranslator-UltraFast-Exp
English | 中文
Experimental. For actual translation, use JMangaTranslator-Fast instead.
JMangaTranslator-UltraFast-Exp translates Japanese manga speech bubbles into Simplified Chinese, one bubble at a time. It is a 229M-parameter non-autoregressive model (a Directed Acyclic Transformer, DAT): it produces the whole translation in one forward pass. On an RTX 3060 it takes 4.5 ms per bubble, about twice as fast as JMangaTranslator-Fast v1. Its quality, however, is clearly lower than v1's.
We publish it as a negative result: the speed is real, but the quality gap did not close. The model, the measurements and the open problem below may be useful to anyone working on non-autoregressive translation.
Architecture, training, evaluation and speed details are in the technical report.
What we learned
- Faster, but worse. On 3,817 Manga109-s bubbles, v1 beats this model by 0.029 COMET on ground-truth text and 0.038 on manga-ocr output. Every paired 95% bootstrap interval excludes zero. In exchange, this model saves about 4 ms per bubble on an RTX 3060 (4.5 ms vs 8.6 ms). v1 is already fast enough to translate each bubble as soon as it is recognized, so we consider the trade not worth it.
- Manga fine-tuning helped only a little. The base model was trained from scratch on light-novel data. Fine-tuning it on 250k manga bubbles raised COMET by 0.004 on both input conditions, a significant gain, but small next to the gap to v1.
- Expensive to train. The base model took 30.1 hours on an RTX 5090. JMangaTranslator-Fast v1, which translates better, took about 15 hours on an RTX 4090 in total. A DAT expands every training example into a graph four times as long as the source, so each training step does much more work than in an ordinary encoder–decoder model.
- The better translation is often already in the graph. A DAT predicts a graph of candidate tokens and then picks one path through it. On two other test sets (Manga200 and Luna1k), we took the 12 beam-search paths from the same forward pass and chose the best one with the reference translation. This oracle choice was 0.043 COMET better than the default path on both sets. Without the reference, the best reranker we tried recovered only +0.020 and +0.0055. These numbers are for the base model before fine-tuning. Picking a better path without a reference, at low cost, is the open problem.
Quality
COMET-22 on 3,817 Manga109-s bubbles, the same set and protocol as the JMangaTranslator-Fast v1 results. The set keeps every bubble on which an OCR model made a mistake, so it is harder than average. "Ground truth" is the annotated text; "manga-ocr" is the same bubbles as read by manga-ocr. Reference translations were produced by Claude.
| System | Parameters | Ground truth | manga-ocr |
|---|---|---|---|
| JMangaTranslator-Fast v1 | 368M | 0.8879 | 0.8315 |
| GalTransl-v4-4B | 4B | 0.8617 | 0.8022 |
| JMangaTranslator-UltraFast-Exp | 229M | 0.8588 | 0.7932 |
| (the same model before manga fine-tuning) | 229M | 0.8543 | 0.7890 |
| NanoSakura-2.2-0.2B | 0.2B | 0.8456 | 0.7827 |
| Sakura-1.5B-Qwen2.5-v1.0 | 1.5B | 0.8392 | 0.7787 |
| Hy-MT2-1.8B-JP-Manga-Finetune-zh-Hans-v1 | 1.8B | 0.8349 | 0.7850 |
| Hy-MT2-1.8B | 1.8B | 0.8299 | 0.7818 |
| opus-mt-ja-zh | 77M | 0.6849 | 0.6532 |
| M2M100-418M | 418M | 0.6582 | 0.6278 |
| NLLB-200-distilled-600M | 600M | 0.6122 | 0.5904 |
Speed
RTX 3060 12 GB, batch size 1, 500 Manga109-s bubbles, from tokenization to the decoded string. Load time includes building the CUDA graphs.
| Model | Backend | p50 | p90 | Load time |
|---|---|---|---|---|
| UltraFast-Exp | CUDA Graphs, fp16 (default) | 4.5 ms | 5.8 ms | 7.9 s |
| UltraFast-Exp | CUDA Graphs + torch.compile, fp16 |
2.8 ms | 3.9 ms | several minutes (260 s in our test) |
| UltraFast-Exp | PyTorch eager, fp32 | 37.4 ms | 38.6 ms | 6.0 s |
| JMangaTranslator-Fast v1 | CUDA Graphs, fp16 | 8.6 ms | 12.8 ms | 3.4 s |
The compiled backend fuses many small GPU kernels and saves another 1.7 ms per bubble, but it must compile for several minutes at every first start. It is therefore off by default.
Usage
The weights and code are on Hugging Face and ModelScope. GitHub holds the code only.
hf download muscgab/JMangaTranslator-UltraFast-Exp --local-dir JMangaTranslator-UltraFast-Exp
# or: modelscope download --model muscgab/JMangaTranslator-UltraFast-Exp --local_dir JMangaTranslator-UltraFast-Exp
cd JMangaTranslator-UltraFast-Exp
pip install -r requirements.txt
python translate.py "堪忍袋の緒が切れた!"
python translate.py < bubbles.txt > translations.txt
python translate.py --compile < bubbles.txt > translations.txt # NVIDIA only; compiles for several minutes first
from jmt_ultrafast import load
tr = load("path/to/JMangaTranslator-UltraFast-Exp") # add compile=True for the compiled backend
print(tr.translate("堪忍袋の緒が切れた!")) # 忍耐断了! (v1: 忍无可忍了!)
Give one bubble per line, with line breaks inside a bubble removed. On an NVIDIA GPU the CUDA Graphs backend is used. Elsewhere the model runs in PyTorch eager mode, which is much slower.
The CPU path of the packaged code was checked bubble by bubble against the measured outputs. The CUDA Graphs and compiled paths have not yet been re-run in this packaged form; the speed numbers above come from the research version of the same code.
Limitations
- Quality. Lower than v1 on both input conditions (see above).
- Typical errors. On the 3,817 ground-truth bubbles, 18 outputs still contain Japanese kana and 7 contain the replacement character U+FFFD. v1 has none of either.
- No context. Each bubble is translated on its own, so names, pronouns and tone may differ between bubbles.
- Machine-made training targets. No human translations were used in training.
Acknowledgments
- OpenSakura. The base model, including its tokenizer, was trained on OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve.
- DA-Transformer. The model follows Huang et al., Directed Acyclic Transformer for Non-Autoregressive Machine Translation (ICML 2022), including lookahead decoding.
- GLAT. Training used glancing training from Qian et al., Glancing Transformer for Non-Autoregressive Neural Machine Translation (ACL 2021).
- FA-DAT. The FA-DAT implementation by Ma et al. (ICLR 2023) was our reference when writing the training objective.
- manga-ocr, whose output forms one of the evaluation conditions.
- Manga109-s, used for evaluation.
License
- Model weights: CC BY-NC-SA 4.0 (
LICENSE-weights.md). Attribution is required, commercial use is not permitted, and adapted models must be shared under the same license. - Code: MIT (
LICENSE).
Training used OpenSakura data (license "other", intended for research and model development) and the author's private manga text. Manga109-s was used for evaluation only. No weights or outputs of the other systems in the comparison are included.
- Downloads last month
- 10