IndicTrans2 1B (en→indic) — ONNX bundle [Q4F16 (4-bit Block Quantization, Lossy)]
This model is part of a suite of optimized/quantized ONNX versions of the base model. Other variants in this direction:
- FP32 (Full Precision / Base):
hari31416/indictrans2-en-indic-1B-ONNX- FP16 (Half Precision):
hari31416/indictrans2-en-indic-1B-ONNX-fp16- INT8 (Dynamic Quantization):
hari31416/indictrans2-en-indic-1B-ONNX-int8- Q4F16 (4-bit Block Quantization):
hari31416/indictrans2-en-indic-1B-ONNX-q4f16(Current)
ONNX-exported and quantized version of ai4bharat/indictrans2-en-indic-1B
for in-browser and local edge inference.
- Precision: Q4F16 (4-bit Block Quantization, Lossy)
- Description: 4-bit quantization with float16 scale factors and block size of 32. Reduces model size significantly.
- Source Pipeline & Details: For pipeline details, benchmarks, and usage instructions, see the indictrans2-onnx-export GitHub repository.
Built for use with Transformers.js and onnxruntime-web in the browser, with fast BPE tokenizer.json files that don't require the SentencePiece WASM runtime.
Performance Visualizations
These charts show overall tradeoffs, language-level parity, and category breakdown.
Performance Tradeoffs & Size Comparison
Compared against the FP32 ONNX oracle on the golden evaluation fixtures.
| Format | Model Size | Exact Match (Token) | Exact Match (Text) | SacreBLEU (Raw) | Latency (Mean) | Speedup vs. FP32 |
|---|---|---|---|---|---|---|
| FP32 | 6.64 GB | 100.00% | 100.00% | 100.00 | 69.5 ms | 1.000x |
| FP16 | 3.32 GB | 99.73% | 99.73% | 100.00 | 74.3 ms | 0.935x |
| INT8 | 1.66 GB | 89.55% | 89.55% | 96.27 | 31.4 ms | 2.125x |
| Q4F16 | 850.5 MB | 82.45% | 82.55% | 91.99 | 58.4 ms | 1.186x |
Language-Level Parity (Q4F16)
Exact match rates and translation quality (SacreBLEU / chrF) per language pair under this precision:
| Language Code | Total Fixtures | Token Match Rate | Text Match Rate | SacreBLEU | SacreBLEU (chrF) |
|---|---|---|---|---|---|
| asm_Beng | 50 | 82.0% | 82.0% | 92.61 | 97.18 |
| ben_Beng | 50 | 82.0% | 82.0% | 90.59 | 95.00 |
| brx_Deva | 50 | 84.0% | 84.0% | 93.64 | 97.53 |
| doi_Deva | 50 | 78.0% | 78.0% | 90.60 | 95.01 |
| gom_Deva | 50 | 76.0% | 76.0% | 87.97 | 95.04 |
| guj_Gujr | 50 | 86.0% | 86.0% | 94.93 | 98.11 |
| hin_Deva | 50 | 92.0% | 92.0% | 97.98 | 98.79 |
| kan_Knda | 50 | 82.0% | 82.0% | 90.60 | 96.27 |
| kas_Arab | 50 | 86.0% | 86.0% | 93.50 | 97.01 |
| mai_Deva | 50 | 90.0% | 90.0% | 92.18 | 97.32 |
| mal_Mlym | 50 | 90.0% | 90.0% | 95.25 | 99.01 |
| mar_Deva | 50 | 82.0% | 82.0% | 91.82 | 97.21 |
| mni_Beng | 50 | 82.0% | 82.0% | 88.86 | 94.03 |
| npi_Deva | 50 | 82.0% | 82.0% | 91.65 | 96.73 |
| ory_Orya | 50 | 90.0% | 90.0% | 95.40 | 98.37 |
| pan_Guru | 50 | 82.0% | 82.0% | 91.51 | 95.60 |
| san_Deva | 50 | 70.0% | 70.0% | 83.93 | 94.57 |
| sat_Olck | 50 | 44.0% | 46.0% | 76.72 | 86.18 |
| snd_Arab | 50 | 82.0% | 82.0% | 92.16 | 96.54 |
| tam_Taml | 50 | 82.0% | 82.0% | 91.81 | 96.57 |
| tel_Telu | 50 | 92.0% | 92.0% | 96.44 | 99.30 |
| urd_Arab | 50 | 98.0% | 98.0% | 99.61 | 99.62 |
Category-Level Parity (Q4F16)
Exact match rates and translation quality grouped by category types:
| Category | Total Fixtures | Token Match Rate | Text Match Rate | SacreBLEU | SacreBLEU (chrF) |
|---|---|---|---|---|---|
| Generic | 286 | 81.1% | 81.1% | 91.48 | 96.25 |
| Lexicon | 264 | 79.5% | 79.5% | 91.13 | 96.07 |
| Numerals | 264 | 84.8% | 85.2% | 92.92 | 96.73 |
| Politics | 286 | 84.3% | 84.3% | 91.88 | 96.28 |
Translation Mismatch Examples
Here is a sample of up to 5 translation mismatches compared to the FP32 oracle. Many mismatches represent minor synonym differences or spacing variations.
Mismatch #1 (Category: Politics)
- Source (eng_Latn → asm_Beng):
What is the capital of India? - Expected (FP32):
ভাৰতৰ ৰাজধানী কি? - Actual (Q4F16):
ভাৰতৰ ৰাজধানী কোনখন?
Mismatch #2 (Category: Numerals)
- Source (eng_Latn → asm_Beng):
Can you please help me find the nearest hospital? - Expected (FP32):
আপুনি অনুগ্ৰহ কৰি মোক নিকটতম চিকিৎসালয়খন বিচাৰি উলিওৱাত সহায় কৰিব পাৰিবনে? - Actual (Q4F16):
আপুনি অনুগ্ৰহ কৰি নিকটতম চিকিৎসালয়খন বিচাৰি উলিওৱাত মোক সহায় কৰিব পাৰিবনে?
Mismatch #3 (Category: Generic)
- Source (eng_Latn → asm_Beng):
The train is delayed by two hours. - Expected (FP32):
ৰে "লখনে দুঘণ্টা পলম কৰে। - Actual (Q4F16):
ৰে'ল দুঘণ্টা পলম হয়।
Mismatch #4 (Category: Generic)
- Source (eng_Latn → asm_Beng):
The crop yield has improved due to good rainfall. - Expected (FP32):
ভাল বৰষুণৰ ফলত শস্যৰ উৎপাদন উন্নত হৈছে। - Actual (Q4F16):
ভাল বৰষুণৰ ফলত শস্যৰ উৎপাদন বৃদ্ধি পাইছে।
Mismatch #5 (Category: Generic)
- Source (eng_Latn → asm_Beng):
A warm cup of tea is perfect for a cold morning. - Expected (FP32):
এক গৰম কাপ চাহ ঠাণ্ডা ৰাতিপুৱাৰ বাবে উপযুক্ত। - Actual (Q4F16):
ঠাণ্ডা ৰাতিপুৱাৰ বাবে এক গৰম কাপ চাহ উপযুক্ত।
Files
encoder_model.onnx(and optional.onnx.dataweights sidecar)decoder_model.onnxanddecoder_with_past_model.onnx(sharedecoder_shared.onnx.datawhen present)translate.py— self-contained Python inference helper (see Usage below)- Fast tokenizer config files (
tokenizer_src.json,tokenizer_tgt.json,tokenizer_meta.json) - Model configuration configs (
config.json,generation_config.json)
Usage Example (Python, onnxruntime)
# translate.py is included in this repo alongside the ONNX bundle.
# You can also find it (and read the full source) at:
# https://github.com/Hari31416/indictrans2-onnx-export/blob/main/src/translate.py
from translate import IndicTransONNX
# Pass a HF repo ID for automatic download, or a local bundle directory path
model = IndicTransONNX("hari31416/indictrans2-en-indic-1B-ONNX-q4f16")
print(model.translate("Who will win the election?", src_lang="eng_Latn", tgt_lang="hin_Deva"))
Required packages:
pip install onnxruntime tokenizers huggingface-hub
License
MIT (preserved from upstream AI4Bharat).
- Downloads last month
- 37
Model tree for hari31416/indictrans2-en-indic-1B-ONNX-q4f16
Base model
ai4bharat/indictrans2-en-indic-1B

