CoolFace
Modelpublic

klebster/g2p_multilingual_byT5_tiny_onnx

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes32downloads
Model Card

Multilingual G2P ByT5 Tiny — ONNX

ONNX export of charsiu/g2p_multilingual_byT5_tiny_16_layers_100. Converts written words to IPA transcriptions across 100 languages.

ArchitectureByT5-tiny (T5ForConditionalGeneration), 20.6M params
PER / WER8.1% / 25.3% (100 langs, 500 words each, greedy)
ONNX FP32 size106 MB (3 graphs: encoder + decoder + decoderwithpast)
ONNX INT8 size27 MB (74% reduction, +0.0005 PER)
Best latency5.30 ms/word (INT8, threads=8) — 1.5x faster than PyTorch CPU
LicenseCC BY 4.0

Quick Start

python
import onnxruntime as ort
from optimum.onnxruntime import ORTModelForSeq2SeqLM
from transformers import AutoTokenizer

so = ort.SessionOptions()
so.intra_op_num_threads = 8   # ~cores/4
so.inter_op_num_threads = 1

model = ORTModelForSeq2SeqLM.from_pretrained(
    "klebster/g2p_multilingual_byT5_tiny_onnx",
    provider="CPUExecutionProvider",
    session_options=so,
)
tokenizer = AutoTokenizer.from_pretrained("klebster/g2p_multilingual_byT5_tiny_onnx")

inputs = tokenizer("<eng-us>: hello", padding=True, add_special_tokens=False, return_tensors="pt")
preds = model.generate(**inputs, num_beams=1, max_length=50)
print(tokenizer.decode(preds[0], skip_special_tokens=True))
# Output: ˈhɛɫoʊ

Input format: <language_code>: word (e.g. <fra>: bonjour, <ger>: Straße). See CharsiuG2P for all 100 language codes.

Benchmark Summary

Tested on 10 languages, 300 words, greedy decoding. Hardware: Intel i9-13900KS, 128 GB DDR5.

Configurationms/wordvs PyTorch CPU
ONNX INT8 + threads=85.301.48x faster
ONNX FP32 + threads=87.471.05x faster
PyTorch CPU (baseline)7.831.00x
ONNX default settings25.090.31x

Key findings: Thread tuning is critical (default ONNX is 3x slower than PyTorch). INT8 quantization + threads=8 is the best config. GPU provides no benefit for this model size. Beam search provides negligible improvement (PER 0.0812 → 0.0808) at 3–8x latency cost.

Correctness: ONNX FP32 output is bit-identical to PyTorch. INT8 degrades PER by only 0.0005.

Evaluation

100 languages, 500 words each (50,000 total), CharsiuG2P test set:

MetricPyTorch CPUONNX FP32ONNX INT8
PER0.08120.08120.0817
WER0.25290.25290.2537

<details> <summary>PER by language (click to expand)</summary>

Language ISO-639-3 + Location / ScriptPyTorch CPUONNX FP32ONNX INT8
ady0.05210.05210.0518
afr0.02510.02510.0243
amh0.33470.33470.3356
ang0.05590.05590.0559
ara0.67100.67100.6716
arg0.01810.01810.0183
arm-e0.02150.02150.0220
arm-w0.00750.00750.0070
aze0.00080.00080.0008
bak0.01240.01240.0138
bel0.00480.00480.0040
bos0.01840.01840.0171
bul0.01580.01580.0161
bur0.17180.17180.1712
cat0.15660.15660.1598
cze0.00770.00770.0079
dan0.21180.21180.2157
dut0.03470.03470.0345
egy0.37520.37520.3781
eng-uk0.15420.15420.1534
eng-us0.21820.21820.2193
enm0.19880.19880.1979
epo0.00020.00020.0000
est0.00720.00720.0070
eus0.00560.00560.0058
fas0.15030.15030.1507
fin0.00110.00110.0011
fra0.00750.00750.0075
fra-qu0.00410.00410.0041
geo0.02540.02540.0258
ger0.04540.04540.0442
gle0.17750.17750.1802
glg0.06900.06900.0687
grc0.29440.29440.2942
gre0.02320.02320.0229
hbs-cyrl0.09670.09670.0972
hbs-latn0.09000.09000.0913
hin0.04210.04210.0417
hun0.02270.02270.0224
ice0.02810.02810.0273
ido0.04840.04840.0479
ina0.05290.05290.0532
ind0.02190.02190.0231
isl0.03410.03410.0341
ita0.02730.02730.0282
jpn0.10500.10500.1062
kaz0.00550.00550.0061
khm0.27870.27870.2850
kor0.07760.07760.0776
kur0.00930.00930.0099
lat-clas0.00940.00940.0092
lat-eccl0.00700.00700.0074
lit0.03160.03160.0316
ltz0.12880.12880.1307
mac0.01130.01130.0110
mlt0.02260.02260.0228
mri0.22140.22140.2198
msa0.00080.00080.0008
nan0.10690.10690.1082
nob0.08740.08740.0888
ori0.00280.00280.0026
pap0.00190.00190.0019
pol0.00570.00570.0057
por-bz0.08660.08660.0875
por-po0.08720.08720.0877
ron0.00720.00720.0072
rus0.02110.02110.0202
san0.07620.07620.0755
slk0.02890.02890.0283
slo0.04150.04150.0418
slv0.18290.18290.1838
sme0.02070.02070.0214
snd0.26360.26360.2685
spa0.00170.00170.0019
spa-latin0.00050.00050.0007
spa-me0.00430.00430.0047
sqi0.01480.01480.0192
srp0.06200.06200.0638
swa0.00250.00250.0025
swe0.00840.00840.0079
syc0.37650.37650.3776
tam0.00990.00990.0112
tat0.00260.00260.0033
tgl0.10660.10660.1066
tha0.09210.09210.0929
tts0.01960.01960.0199
tuk0.01190.01190.0123
tur0.00060.00060.0004
uig0.03350.03350.0340
ukr0.03380.03380.0332
urd0.23520.23520.2374
uzb0.00430.00430.0045
vie-c0.01440.01440.0144
vie-n0.00750.00750.0076
vie-s0.01120.01120.0107
wel-nw0.09670.09670.0990
wel-sw0.12910.12910.1321
yue0.22380.22380.2242
zho-s0.29860.29860.2992
zho-t0.34590.34590.3478
AVERAGE (arithmetic mean)0.08120.08120.0817

</details>

<details> <summary>WER by language (click to expand)</summary>

Language ISO-639-3 + Location / ScriptPyTorch CPUONNX FP32ONNX INT8
ady0.27000.27000.2680
afr0.12400.12400.1220
amh0.99400.99400.9940
ang0.32600.32600.3200
ara1.00001.00001.0000
arg0.09600.09600.1000
arm-e0.11200.11200.1140
arm-w0.05200.05200.0480
aze0.00600.00600.0060
bak0.09200.09200.1020
bel0.05400.05400.0440
bos0.06200.06200.0600
bul0.10800.10800.1140
bur0.50800.50800.5000
cat0.60000.60000.5980
cze0.04200.04200.0440
dan0.67400.67400.6700
dut0.19400.19400.1920
egy0.58200.58200.5920
eng-uk0.31800.31800.3260
eng-us0.48400.48400.4860
enm0.71600.71600.7120
epo0.00200.00200.0000
est0.06400.06400.0620
eus0.02200.02200.0220
fas0.56400.56400.5660
fin0.00800.00800.0080
fra0.02000.02000.0200
fra-qu0.02000.02000.0200
geo0.23600.23600.2400
ger0.17200.17200.1700
gle0.66600.66600.6720
glg0.35800.35800.3560
grc0.73400.73400.7360
gre0.16200.16200.1600
hbs-cyrl0.51800.51800.5200
hbs-latn0.46200.46200.4680
hin0.21400.21400.2060
hun0.11400.11400.1140
ice0.17000.17000.1680
ido0.23000.23000.2260
ina0.20400.20400.2020
ind0.10200.10200.1040
isl0.22600.22600.2260
ita0.20000.20000.2040
jpn0.23800.23800.2400
kaz0.04200.04200.0420
khm0.73200.73200.7220
kor0.09400.09400.0940
kur0.05200.05200.0480
lat-clas0.05800.05800.0560
lat-eccl0.03800.03800.0400
lit0.18600.18600.1900
ltz0.38000.38000.3840
mac0.08800.08800.0860
mlt0.08800.08800.0900
mri0.67400.67400.6740
msa0.00400.00400.0040
nan0.41000.41000.4140
nob0.34000.34000.3340
ori0.02200.02200.0200
pap0.01200.01200.0120
pol0.02400.02400.0240
por-bz0.43200.43200.4320
por-po0.39200.39200.3980
ron0.05000.05000.0500
rus0.12600.12600.1220
san0.44000.44000.4360
slk0.11400.11400.1160
slo0.21800.21800.2200
slv0.58600.58600.5880
sme0.13600.13600.1380
snd0.60000.60000.6120
spa0.01600.01600.0180
spa-latin0.00400.00400.0060
spa-me0.03800.03800.0400
sqi0.07400.07400.0900
srp0.44800.44800.4440
swa0.01800.01800.0180
swe0.06000.06000.0560
syc0.89800.89800.9080
tam0.05000.05000.0560
tat0.01400.01400.0180
tgl0.44600.44600.4500
tha0.21400.21400.2160
tts0.07800.07800.0760
tuk0.11000.11000.1140
tur0.00400.00400.0020
uig0.14800.14800.1540
ukr0.22400.22400.2180
urd0.58600.58600.5900
uzb0.02200.02200.0240
vie-c0.04000.04000.0400
vie-n0.03000.03000.0300
vie-s0.04000.04000.0380
wel-nw0.36600.36600.3720
wel-sw0.41200.41200.4260
yue0.40400.40400.4080
zho-s0.53200.53200.5300
zho-t0.56000.56000.5620
AVERAGE (arithmetic mean)0.25290.25290.2537

</details>

Links

Citation

If you use this ONNX model, please cite both the original paper and the ONNX export:

bibtex
@misc{zhu2022byt5modelmassivelymultilingual,
      title={ByT5 model for massively multilingual grapheme-to-phoneme conversion},
      author={Jian Zhu and Cong Zhang and David Jurgens},
      year={2022},
      eprint={2204.03067},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2204.03067},
}
bibtex
@misc{noel2025g2pmultilingualbyT5tinyonnx,
      title={Multilingual G2P ByT5 Tiny — ONNX export},
      author={Kleber Noel},
      year={2026},
      month={apr},
      note={Published 2026-04-07},
      url={https://huggingface.co/klebster/g2p_multilingual_byT5_tiny_onnx},
}

Known Issues

German IPA quality: non-standard dialect

The model does not reliably produce Standard German (Hochdeutsch). Observed issues include use of alveolar flap /ɾ/ where Standard German uses uvular fricative /ʁ/, among other systematic deviations. See CharsiuG2P issue #20.

Spanish dialect dictionaries: spa and spa-me are identical

The spa (European Spanish) and spa-me (Mexican Spanish) dictionaries are identical in the upstream CharsiuG2P repository. They should differ in the /s/–/θ/ distinction (ceceo/seseo). See CharsiuG2P issue #15.