IndicTrans2 1B (en→indic) — ONNX bundle [Q4F16 (4-bit Block Quantization, Lossy)]

This model is part of a suite of optimized/quantized ONNX versions of the base model. Other variants in this direction:

ONNX-exported and quantized version of ai4bharat/indictrans2-en-indic-1B for in-browser and local edge inference.

  • Precision: Q4F16 (4-bit Block Quantization, Lossy)
  • Description: 4-bit quantization with float16 scale factors and block size of 32. Reduces model size significantly.
  • Source Pipeline & Details: For pipeline details, benchmarks, and usage instructions, see the indictrans2-onnx-export GitHub repository.

Built for use with Transformers.js and onnxruntime-web in the browser, with fast BPE tokenizer.json files that don't require the SentencePiece WASM runtime.

Performance Visualizations

These charts show overall tradeoffs, language-level parity, and category breakdown.

Overall Tradeoffs Language-Level Parity Category breakdown

Performance Tradeoffs & Size Comparison

Compared against the FP32 ONNX oracle on the golden evaluation fixtures.

Format Model Size Exact Match (Token) Exact Match (Text) SacreBLEU (Raw) Latency (Mean) Speedup vs. FP32
FP32 6.64 GB 100.00% 100.00% 100.00 69.5 ms 1.000x
FP16 3.32 GB 99.73% 99.73% 100.00 74.3 ms 0.935x
INT8 1.66 GB 89.55% 89.55% 96.27 31.4 ms 2.125x
Q4F16 850.5 MB 82.45% 82.55% 91.99 58.4 ms 1.186x

Language-Level Parity (Q4F16)

Exact match rates and translation quality (SacreBLEU / chrF) per language pair under this precision:

Language Code Total Fixtures Token Match Rate Text Match Rate SacreBLEU SacreBLEU (chrF)
asm_Beng 50 82.0% 82.0% 92.61 97.18
ben_Beng 50 82.0% 82.0% 90.59 95.00
brx_Deva 50 84.0% 84.0% 93.64 97.53
doi_Deva 50 78.0% 78.0% 90.60 95.01
gom_Deva 50 76.0% 76.0% 87.97 95.04
guj_Gujr 50 86.0% 86.0% 94.93 98.11
hin_Deva 50 92.0% 92.0% 97.98 98.79
kan_Knda 50 82.0% 82.0% 90.60 96.27
kas_Arab 50 86.0% 86.0% 93.50 97.01
mai_Deva 50 90.0% 90.0% 92.18 97.32
mal_Mlym 50 90.0% 90.0% 95.25 99.01
mar_Deva 50 82.0% 82.0% 91.82 97.21
mni_Beng 50 82.0% 82.0% 88.86 94.03
npi_Deva 50 82.0% 82.0% 91.65 96.73
ory_Orya 50 90.0% 90.0% 95.40 98.37
pan_Guru 50 82.0% 82.0% 91.51 95.60
san_Deva 50 70.0% 70.0% 83.93 94.57
sat_Olck 50 44.0% 46.0% 76.72 86.18
snd_Arab 50 82.0% 82.0% 92.16 96.54
tam_Taml 50 82.0% 82.0% 91.81 96.57
tel_Telu 50 92.0% 92.0% 96.44 99.30
urd_Arab 50 98.0% 98.0% 99.61 99.62

Category-Level Parity (Q4F16)

Exact match rates and translation quality grouped by category types:

Category Total Fixtures Token Match Rate Text Match Rate SacreBLEU SacreBLEU (chrF)
Generic 286 81.1% 81.1% 91.48 96.25
Lexicon 264 79.5% 79.5% 91.13 96.07
Numerals 264 84.8% 85.2% 92.92 96.73
Politics 286 84.3% 84.3% 91.88 96.28

Translation Mismatch Examples

Here is a sample of up to 5 translation mismatches compared to the FP32 oracle. Many mismatches represent minor synonym differences or spacing variations.

Mismatch #1 (Category: Politics)

  • Source (eng_Latn → asm_Beng): What is the capital of India?
  • Expected (FP32): ভাৰতৰ ৰাজধানী কি?
  • Actual (Q4F16): ভাৰতৰ ৰাজধানী কোনখন?

Mismatch #2 (Category: Numerals)

  • Source (eng_Latn → asm_Beng): Can you please help me find the nearest hospital?
  • Expected (FP32): আপুনি অনুগ্ৰহ কৰি মোক নিকটতম চিকিৎসালয়খন বিচাৰি উলিওৱাত সহায় কৰিব পাৰিবনে?
  • Actual (Q4F16): আপুনি অনুগ্ৰহ কৰি নিকটতম চিকিৎসালয়খন বিচাৰি উলিওৱাত মোক সহায় কৰিব পাৰিবনে?

Mismatch #3 (Category: Generic)

  • Source (eng_Latn → asm_Beng): The train is delayed by two hours.
  • Expected (FP32): ৰে "লখনে দুঘণ্টা পলম কৰে।
  • Actual (Q4F16): ৰে'ল দুঘণ্টা পলম হয়।

Mismatch #4 (Category: Generic)

  • Source (eng_Latn → asm_Beng): The crop yield has improved due to good rainfall.
  • Expected (FP32): ভাল বৰষুণৰ ফলত শস্যৰ উৎপাদন উন্নত হৈছে।
  • Actual (Q4F16): ভাল বৰষুণৰ ফলত শস্যৰ উৎপাদন বৃদ্ধি পাইছে।

Mismatch #5 (Category: Generic)

  • Source (eng_Latn → asm_Beng): A warm cup of tea is perfect for a cold morning.
  • Expected (FP32): এক গৰম কাপ চাহ ঠাণ্ডা ৰাতিপুৱাৰ বাবে উপযুক্ত।
  • Actual (Q4F16): ঠাণ্ডা ৰাতিপুৱাৰ বাবে এক গৰম কাপ চাহ উপযুক্ত।

Files

  • encoder_model.onnx (and optional .onnx.data weights sidecar)
  • decoder_model.onnx and decoder_with_past_model.onnx (share decoder_shared.onnx.data when present)
  • translate.py — self-contained Python inference helper (see Usage below)
  • Fast tokenizer config files (tokenizer_src.json, tokenizer_tgt.json, tokenizer_meta.json)
  • Model configuration configs (config.json, generation_config.json)

Usage Example (Python, onnxruntime)

# translate.py is included in this repo alongside the ONNX bundle.
# You can also find it (and read the full source) at:
#   https://github.com/Hari31416/indictrans2-onnx-export/blob/main/src/translate.py

from translate import IndicTransONNX

# Pass a HF repo ID for automatic download, or a local bundle directory path
model = IndicTransONNX("hari31416/indictrans2-en-indic-1B-ONNX-q4f16")
print(model.translate("Who will win the election?", src_lang="eng_Latn", tgt_lang="hin_Deva"))

Required packages:

pip install onnxruntime tokenizers huggingface-hub

License

MIT (preserved from upstream AI4Bharat).

Downloads last month
37
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hari31416/indictrans2-en-indic-1B-ONNX-q4f16

Quantized
(4)
this model

Collection including hari31416/indictrans2-en-indic-1B-ONNX-q4f16