CoolFace
Modelpublic

hari31416/indictrans2-indic-indic-dist-320M-ONNX-int8

sourceHugging Facemitupdated 2mo agoView on Hugging Face
1likes125downloads
Model Card

IndicTrans2 200M (indic→indic) — ONNX bundle [INT8 (Dynamic Quantization)]

[!TIP] This model is part of a suite of optimized/quantized ONNX versions of the base model. Other variants in this direction: - FP32 (Full Precision / Base): `hari31416/indictrans2-indic-indic-dist-320M-ONNX` - FP16 (Half Precision): `hari31416/indictrans2-indic-indic-dist-320M-ONNX-fp16` - INT8 (Dynamic Quantization): `hari31416/indictrans2-indic-indic-dist-320M-ONNX-int8` (Current) - Q4F16 (4-bit Block Quantization): `hari31416/indictrans2-indic-indic-dist-320M-ONNX-q4f16`

ONNX-exported and quantized version of `ai4bharat/indictrans2-indic-indic-dist-320M` for in-browser and local edge inference.

  • —Precision: INT8 (Dynamic Quantization)
  • —Description: Dynamic INT8 quantization of the encoder and decoder. Highly recommended for CPU environments.
  • —Source Pipeline & Details: For pipeline details, benchmarks, and usage instructions, see the indictrans2-onnx-export GitHub repository.

Built for use with Transformers.js and onnxruntime-web in the browser, with fast BPE tokenizer.json files that don't require the SentencePiece WASM runtime.

Performance Visualizations

These charts show overall tradeoffs, language-level parity, and category breakdown.

[image] [image] [image]

Performance Tradeoffs & Size Comparison

Compared against the FP32 ONNX oracle on the golden evaluation fixtures.

FormatModel SizeExact Match (Token)Exact Match (Text)SacreBLEU (Raw)Latency (Mean)Speedup vs. FP32
FP321.87 GB100.00%100.00%100.0023.0 ms1.000x
FP16980.2 MB99.82%99.82%100.0027.4 ms0.840x
INT8497.1 MB72.18%72.36%87.1316.5 ms1.475x
Q4F16711.6 MB45.91%46.36%71.6428.3 ms0.829x

Language-Level Parity (INT8)

Exact match rates and translation quality (SacreBLEU / chrF) per language pair under this precision:

Language CodeTotal FixturesToken Match RateText Match RateSacreBLEUSacreBLEU (chrF)
asm_Beng5076.0%76.0%87.3295.24
ben_Beng5068.0%68.0%88.3495.76
brx_Deva5066.0%66.0%82.5493.75
doi_Deva5076.0%76.0%90.1194.95
gom_Deva5076.0%76.0%86.9593.38
guj_Gujr5088.0%88.0%96.1698.54
hin_Deva5080.0%80.0%91.6496.19
kan_Knda5084.0%84.0%90.2096.61
kas_Arab5070.0%70.0%85.5392.09
mai_Deva5082.0%84.0%92.4795.03
mal_Mlym5078.0%78.0%88.3694.82
mar_Deva5082.0%82.0%91.5195.37
mni_Beng5068.0%68.0%86.4193.50
npi_Deva5058.0%60.0%73.6088.80
ory_Orya5068.0%68.0%85.3195.18
pan_Guru5086.0%86.0%94.7296.92
san_Deva5070.0%70.0%79.7091.60
sat_Olck5060.0%60.0%86.3092.15
snd_Arab5032.0%32.0%42.1149.16
tam_Taml5066.0%66.0%82.3193.73
tel_Telu5076.0%76.0%86.9494.19
urd_Arab5078.0%78.0%92.2496.00

Category-Level Parity (INT8)

Exact match rates and translation quality grouped by category types:

CategoryTotal FixturesToken Match RateText Match RateSacreBLEUSacreBLEU (chrF)
Generic28667.8%68.2%82.5491.29
Lexicon26468.9%68.9%84.7090.98
Numerals26473.5%73.5%87.1393.50
Politics28678.3%78.7%88.1593.69

Translation Mismatch Examples

Here is a sample of up to 5 translation mismatches compared to the FP32 oracle. Many mismatches represent minor synonym differences or spacing variations.

Mismatch #1 (Category: Numerals)

  • —Source (asm_Beng → ben_Beng): আপুনি অনুগ্ৰহ কৰি মোক নিকটতম চিকিৎসালয়খন বিচাৰি উলিওৱাত সহায় কৰিব পাৰিবনে?
  • —Expected (FP32): আপনি কি দয়া করে নিকটতম হাসপাতাল খুঁজে পেতে আমার সাহায্য করতে পারেন?
  • —Actual (INT8): আপনি কি দয়া করে নিকটতম হাসপাতাল খুঁজে পেতে আমাকে সাহায্য করতে পারেন?

Mismatch #2 (Category: Lexicon)

  • —Source (asm_Beng → ben_Beng): মই দিল্লীলৈ বিমানৰ টিকট এখন বুক কৰিব বিচাৰো।
  • —Expected (FP32): আমি দিল্লিতে যাওয়ার জন্য একটি বিমানের টিকিট বুক করতে চাই।
  • —Actual (INT8): আমি দিল্লিতে একটি বিমানের টিকিট বুক করতে চাই।

Mismatch #3 (Category: Numerals)

  • —Source (asm_Beng → ben_Beng): মোৰ ফোন নম্বৰটো হৈছে + 91-9876543210।
  • —Expected (FP32): আমার ফোন নম্বর হল + 91-9876543210।
  • —Actual (INT8): আমার ফোন নম্বর + 91-9876543210।

Mismatch #4 (Category: Generic)

  • —Source (asm_Beng → ben_Beng): কৃত্ৰিম বুদ্ধিমত্তাই শিক্ষাৰ ক্ষেত্ৰত পৰিৱৰ্তন কঢ়িয়াই আহিছে।
  • —Expected (FP32): কৃত্রিম বুদ্ধিমত্তা শিক্ষার ক্ষেত্রে পরিবর্তন আনছে।
  • —Actual (INT8): কৃত্রিম বুদ্ধিমত্তা শিক্ষার ক্ষেত্রে পরিবর্তন নিয়ে এসেছে।

Mismatch #5 (Category: Numerals)

  • —Source (asm_Beng → ben_Beng): ৱেবছাইটটোৰ ব্যৱহাৰকাৰী আন্তঃপৃষ্ঠ পৰিষ্কাৰ আৰু আধুনিক।
  • —Expected (FP32): ওয়েবসাইটের ব্যবহারকারী ইন্টারফেস পরিষ্কার এবং আধুনিক।
  • —Actual (INT8): ওয়েবসাইটের ব্যবহারকারী ইন্টারফেসটি পরিষ্কার এবং আধুনিক।

Files

  • —encoder_model.onnx (and optional .onnx.data weights sidecar)
  • —decoder_model.onnx and decoder_with_past_model.onnx (share decoder_shared.onnx.data when present)
  • —translate.py — self-contained Python inference helper (see Usage below)
  • —Fast tokenizer config files (tokenizer_src.json, tokenizer_tgt.json, tokenizer_meta.json)
  • —Model configuration configs (config.json, generation_config.json)

Usage Example (Python, onnxruntime)

python
# translate.py is included in this repo alongside the ONNX bundle.
# You can also find it (and read the full source) at:
#   https://github.com/Hari31416/indictrans2-onnx-export/blob/main/src/translate.py

from translate import IndicTransONNX

# Pass a HF repo ID for automatic download, or a local bundle directory path
model = IndicTransONNX("hari31416/indictrans2-indic-indic-dist-320M-ONNX-int8")
print(model.translate("चुनाव कौन जीतेगा?", src_lang="hin_Deva", tgt_lang="tam_Taml"))

Required packages:

bash
pip install onnxruntime tokenizers huggingface-hub

License

MIT (preserved from upstream AI4Bharat).