calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8
Parakeet TDT 0.6B v3 pt-BR TAGARELA — ONNX INT8 Dynamic
🇧🇷 Quantização INT8 dinâmica do melhor ASR open-source para português brasileiro hoje, pronta para CPU. Compatível direto com Whispering, `onnx-asr` e qualquer pipeline ONNX Runtime. 🇬🇧 Dynamic INT8 quantization of the strongest open-source pt-BR ASR available today, CPU-ready. Drop-in compatible with Whispering, onnx-asr, and any ONNX Runtime pipeline.⚡ Quick Start
pip install "onnx-asr[cpu,hub]"import onnx_asr
model = onnx_asr.load_model(
"nemo-conformer-tdt",
"calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8",
)
print(model.recognize("audio.wav", language="pt"))🇧🇷 Sobre este modelo (pt-BR)
Este repositório contém uma quantização pós-treino INT8 dinâmica do modelo `alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx` — o ASR Parakeet TDT 0.6B v3 da NVIDIA, fine-tuned em pt-BR pelo dataset TAGARELA por Alefiury.
Por que existe:
- Original FP32 ocupa ~2.55 GB em disco e RAM. Pesado para deploy em apps desktop como Whispering.
- INT8 dynamic reduz para ~890 MB (-65 %) sem retreino, sem dataset de calibração.
- MatMul-only quantization (camadas Conv e LayerNorm preservadas em FP) — abordagem conservadora, espera-se < 1 pp de degradação WER.
Para quem é:
- Usuários de Whispering / Epicenter que querem transcrição em português local de qualidade.
- Devs Python que precisam de ASR pt-BR sem GPU, sem cloud, sem custo por minuto.
- Aplicações que rodam em desktop / edge com restrição de RAM e disco.
Quando NÃO usar:
- Você tem GPU e quer máxima qualidade → use o FP32 original.
- Você precisa de transcrição multi-idioma → use o Parakeet 0.6B v3 NVIDIA original.
- Workloads de batch em servidor → use static quantization com calibração própria.
🇬🇧 About this model (English)
INT8 dynamic post-training quantization of `alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx` — NVIDIA's Parakeet TDT 0.6B v3 fine-tuned for Brazilian Portuguese on the TAGARELA dataset by Alefiury.
Why it exists: Original FP32 weights are ~2.55 GB. Too heavy for desktop apps. Dynamic INT8 cuts that to ~890 MB (-65 %) with no retraining and no calibration dataset. Conservative MatMul-only quantization preserves Conv / LayerNorm layers in FP — expected WER degradation < 1 pp.
Source model lineage
nvidia/parakeet-tdt-0.6b-v3 (multilingual base, FP32)
└── alexandreacff/parakeet-tdt-0.6b-v3-ptBR-plus (pt-BR fine-tune)
└── alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx (TAGARELA fine-tune + ONNX export, FP32)
└── this repo (INT8 dynamic quantization)All credit for training, dataset curation, and original weights goes to Alefiury, Alexandre Acff, and the NVIDIA NeMo team. This repository contributes only post-training quantization.
Benchmarks (inherited from base model)
The base TAGARELA model achieves state-of-the-art pt-BR WER among similarly-sized open models. Numbers below are from the FP32 base — INT8 dynamic typically degrades < 1 pp WER on Conformer ASR but was not formally re-evaluated for this checkpoint. Bench it on your own data before production.
Prepared speech (WER ↓)
Spontaneous speech (WER ↓)
⭐ Best-in-class for spontaneous Brazilian Portuguese among similarly-sized open models.
Quantization details
File sizes
Why MatMul-only?
ASR Conformer encoders are dominated by MatMul ops (attention Q/K/V/O projections + feed-forward). Quantizing these captures most of the size win. Conv layers (~10 % of params) and LayerNorm carry less weight but are more sensitive to quantization noise — leaving them in FP is the documented best practice from NVIDIA / Microsoft for Conformer-family models.
Files
config.json # NeMo Conformer-TDT metadata
decoder_joint-model.int8.onnx # TDT decoder + joint network (INT8)
encoder-model.int8.onnx # FastConformer encoder (INT8)
nemo128.onnx # 128-dim mel-spectrogram preprocessor (FP, unchanged)
vocab.txt # SentencePiece tokenizer vocabulary
README.md # this fileUsage
With Whispering (Epicenter)
Whispering uses `transcribe-rs` under the hood, which expects the exact filenames in this repo (encoder-model.int8.onnx, decoder_joint-model.int8.onnx). To install:
Windows:
$dst = "$env:APPDATA\com.bradenwong.whispering\models\parakeet\parakeet-tdt-0.6b-v3-ptBR-int8"
hf download calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8 --local-dir $dstmacOS / Linux:
dst="$HOME/.config/com.bradenwong.whispering/models/parakeet/parakeet-tdt-0.6b-v3-ptBR-int8"
hf download calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8 --local-dir "$dst"Restart Whispering. The model appears in the model selector as parakeet-tdt-0.6b-v3-ptBR-int8.
With onnx-asr (Python)
pip install "onnx-asr[cpu,hub]"import onnx_asr
model = onnx_asr.load_model(
"nemo-conformer-tdt",
"calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8",
)
print(model.recognize("audio.wav", language="pt"))Direct ONNX Runtime
The model follows the standard NVIDIA NeMo Conformer-TDT export layout. Any inference stack that supports nemo-conformer-tdt (e.g. NeMo's own nemo.collections.asr ONNX backend, custom Rust / C++ runtimes) should accept it after quantized op-set negotiation.
Reproducibility
from onnxruntime.quantization import quantize_dynamic, QuantType
# Encoder
quantize_dynamic(
model_input="encoder-model.onnx",
model_output="encoder-model.int8.onnx",
weight_type=QuantType.QInt8,
op_types_to_quantize=["MatMul"],
use_external_data_format=False,
)
# Decoder + joint
quantize_dynamic(
model_input="decoder_joint-model.onnx",
model_output="decoder_joint-model.int8.onnx",
weight_type=QuantType.QInt8,
op_types_to_quantize=["MatMul"],
)Environment: Python 3.11, onnxruntime 1.23.2, onnx 1.17.0, Windows 11.
Limitations & known caveats
- No formal WER benchmark on the quantized weights. Architecture is preserved and INT8 dynamic typically costs < 1 pp WER on Conformer ASR — but verify on your own pt-BR test set before production deployment.
- CPU-targeted. Dynamic INT8 may not accelerate GPU inference (CUDA EP). For GPU, prefer FP16 or static quantization with calibration.
- Single-file ONNX limit (~2 GB protobuf). Encoder fits comfortably at 839 MB; if you re-quantize differently and exceed 2 GB, switch
use_external_data_format=True. - Want smaller? INT4 weight-only via
MatMul4BitsQuantizerreduces the encoder to ~280-380 MB but typically costs 1-3 pp WER and requires runtimes that supportMatMulNBitsop (ONNX Runtime ≥ 1.16). Whispering'stranscribe-rsdoes not currently support INT4.
License
Inherits CC-BY-4.0 from the base model. Attribution is required for any downstream use.
Citation
If you use this quantized model, please cite the original work:
@misc{parakeet-tdt-ptbr-tagarela,
author = {Alefiury},
title = {Parakeet TDT 0.6B v3 pt-BR TAGARELA (ONNX)},
year = {2025},
url = {https://huggingface.co/alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx}
}
@article{xu2023tdt,
title = {Efficient Sequence Transduction by Jointly Predicting Tokens and Durations},
author = {Xu, Hainan and Jia, Fei and Majumdar, Somshubra and others},
year = {2023},
eprint = {2304.06795},
archivePrefix = {arXiv}
}Acknowledgements
- [Alefiury](https://huggingface.co/alefiury) — TAGARELA fine-tuning + ONNX export
- [Alexandre Acff](https://huggingface.co/alexandreacff) — pt-BR base fine-tune
- [NVIDIA NeMo team](https://github.com/NVIDIA/NeMo) — Parakeet architecture, FastConformer encoder, TDT decoder, multilingual pretraining
- [istupakov/onnx-asr](https://github.com/istupakov/onnx-asr) — reference ONNX inference framework
- [EpicenterHQ / Whispering](https://github.com/EpicenterHQ/epicenter) — local-first transcription app this quantization targets
