OPUS-MT ONNX int8 packages

The official Helsinki-NLP OPUS-MT Marian models exported to ONNX and quantized to int8 for on-device translation with ONNX Runtime (used by the Hi Talk Android app, among others). Each opus-mt-<src>-<tgt>/ folder holds one direction:

Files per pair

  • encoder_model.onnx, decoder_model.onnx, decoder_with_past_model.onnx โ€” ๐Ÿค— Optimum export (text2text-generation-with-past) of the original weights, dynamically quantized to int8 with ONNX Runtime (quantize_dynamic).
  • tokenizer.bin โ€” the source SentencePiece unigram model and the Marian vocabulary (source.spm, vocab.json) repacked into one binary table (layout documented in the converter's mtok.py); the originals stay in the Helsinki-NLP repositories.
  • manifest.json โ€” byte sizes and SHA-256 sums of the four files, for consumers that pin and verify what they download.

The graphs bind the input/output names Optimum assigns (input_ids, attention_mask, encoder_hidden_states, past_key_values.N.*, present.N.*, logits) and run with any ONNX Runtime.

Attribution and licence

The models are derivative works of the Helsinki-NLP OPUS-MT models listed above, trained by the OPUS-MT project (Jรถrg Tiedemann and Santhosh Thottingal, OPUS-MT โ€” Building open translation services for the World, EAMT 2020) with the Marian NMT framework, and published under CC-BY-4.0. This repository is likewise published under CC-BY-4.0. Changes made to the originals: ONNX export, int8 dynamic quantization, and repacking of the tokenizer files; the weights are otherwise unchanged.

Converted with helsinki-onnx-converter; the exact pipeline, checks and the tokenizer.bin layout are documented there.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Zweiarm/opus-mt-onnx-int8

Quantized
(5)
this model