CoolFace
Modelpublic

calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
7likes136downloads
Model Card

Parakeet TDT 0.6B v3 pt-BR TAGARELA — ONNX INT8 Dynamic

🇧🇷 Quantização INT8 dinâmica do melhor ASR open-source para português brasileiro hoje, pronta para CPU. Compatível direto com Whispering, `onnx-asr` e qualquer pipeline ONNX Runtime. 🇬🇧 Dynamic INT8 quantization of the strongest open-source pt-BR ASR available today, CPU-ready. Drop-in compatible with Whispering, onnx-asr, and any ONNX Runtime pipeline.

⚡ Quick Start

bash
pip install "onnx-asr[cpu,hub]"
python
import onnx_asr

model = onnx_asr.load_model(
    "nemo-conformer-tdt",
    "calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8",
)
print(model.recognize("audio.wav", language="pt"))

🇧🇷 Sobre este modelo (pt-BR)

Este repositório contém uma quantização pós-treino INT8 dinâmica do modelo `alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx` — o ASR Parakeet TDT 0.6B v3 da NVIDIA, fine-tuned em pt-BR pelo dataset TAGARELA por Alefiury.

Por que existe:

  • —Original FP32 ocupa ~2.55 GB em disco e RAM. Pesado para deploy em apps desktop como Whispering.
  • —INT8 dynamic reduz para ~890 MB (-65 %) sem retreino, sem dataset de calibração.
  • —MatMul-only quantization (camadas Conv e LayerNorm preservadas em FP) — abordagem conservadora, espera-se < 1 pp de degradação WER.

Para quem é:

  • —Usuários de Whispering / Epicenter que querem transcrição em português local de qualidade.
  • —Devs Python que precisam de ASR pt-BR sem GPU, sem cloud, sem custo por minuto.
  • —Aplicações que rodam em desktop / edge com restrição de RAM e disco.

Quando NÃO usar:

  • —Você tem GPU e quer máxima qualidade → use o FP32 original.
  • —Você precisa de transcrição multi-idioma → use o Parakeet 0.6B v3 NVIDIA original.
  • —Workloads de batch em servidor → use static quantization com calibração própria.

🇬🇧 About this model (English)

INT8 dynamic post-training quantization of `alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx` — NVIDIA's Parakeet TDT 0.6B v3 fine-tuned for Brazilian Portuguese on the TAGARELA dataset by Alefiury.

Why it exists: Original FP32 weights are ~2.55 GB. Too heavy for desktop apps. Dynamic INT8 cuts that to ~890 MB (-65 %) with no retraining and no calibration dataset. Conservative MatMul-only quantization preserves Conv / LayerNorm layers in FP — expected WER degradation < 1 pp.


Source model lineage

nvidia/parakeet-tdt-0.6b-v3 (multilingual base, FP32)
   └── alexandreacff/parakeet-tdt-0.6b-v3-ptBR-plus (pt-BR fine-tune)
         └── alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx (TAGARELA fine-tune + ONNX export, FP32)
               └── this repo (INT8 dynamic quantization)

All credit for training, dataset curation, and original weights goes to Alefiury, Alexandre Acff, and the NVIDIA NeMo team. This repository contributes only post-training quantization.


Benchmarks (inherited from base model)

The base TAGARELA model achieves state-of-the-art pt-BR WER among similarly-sized open models. Numbers below are from the FP32 base — INT8 dynamic typically degrades < 1 pp WER on Conformer ASR but was not formally re-evaluated for this checkpoint. Bench it on your own data before production.

Prepared speech (WER ↓)

ModelCETUCCommon VoiceMLS (PT)MTEDx (PT)Avg
Parakeet TDT 0.6B v3 ptBR TAGARELA (FP32)0.0060.0510.1080.1330.075
ElevenLabs Scribe v20.0180.0470.0380.1380.060

Spontaneous speech (WER ↓)

ModelALIPC-ORALNURC-RSP2010NURC-SPMuPeAvg
Parakeet TDT 0.6B v3 ptBR TAGARELA (FP32)0.2130.1370.1380.1040.1600.1200.143

⭐ Best-in-class for spontaneous Brazilian Portuguese among similarly-sized open models.


Quantization details

ItemValue
MethodDynamic post-training quantization (no calibration data needed)
Toolonnxruntime.quantization.quantize_dynamic (onnxruntime 1.23.2)
Weight typeQInt8 (signed 8-bit integer)
Ops quantizedMatMul only — conservative, preserves Conv / LayerNorm in FP
External dataDisabled — single-file ONNX, easy to deploy
Hardware usedIntel i7-13700K (Windows 11), CPU only
Total quantization time~14 seconds

File sizes

FileOriginal (FP32)This repo (INT8)Reduction
Encoder (encoder-model.int8.onnx)~2.48 GB839 MB-66 %
Decoder + joint (decoder_joint-model.int8.onnx)72.5 MB51.1 MB-30 %
Mel preprocessor (nemo128.onnx)137 KB137 KB0 %
Vocab (vocab.txt)92 KB92 KB0 %
Total~2.55 GB~890 MB-65 %

Why MatMul-only?

ASR Conformer encoders are dominated by MatMul ops (attention Q/K/V/O projections + feed-forward). Quantizing these captures most of the size win. Conv layers (~10 % of params) and LayerNorm carry less weight but are more sensitive to quantization noise — leaving them in FP is the documented best practice from NVIDIA / Microsoft for Conformer-family models.


Files

config.json                       # NeMo Conformer-TDT metadata
decoder_joint-model.int8.onnx     # TDT decoder + joint network (INT8)
encoder-model.int8.onnx           # FastConformer encoder (INT8)
nemo128.onnx                      # 128-dim mel-spectrogram preprocessor (FP, unchanged)
vocab.txt                         # SentencePiece tokenizer vocabulary
README.md                         # this file

Usage

With Whispering (Epicenter)

Whispering uses `transcribe-rs` under the hood, which expects the exact filenames in this repo (encoder-model.int8.onnx, decoder_joint-model.int8.onnx). To install:

Windows:

powershell
$dst = "$env:APPDATA\com.bradenwong.whispering\models\parakeet\parakeet-tdt-0.6b-v3-ptBR-int8"
hf download calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8 --local-dir $dst

macOS / Linux:

bash
dst="$HOME/.config/com.bradenwong.whispering/models/parakeet/parakeet-tdt-0.6b-v3-ptBR-int8"
hf download calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8 --local-dir "$dst"

Restart Whispering. The model appears in the model selector as parakeet-tdt-0.6b-v3-ptBR-int8.

With onnx-asr (Python)

bash
pip install "onnx-asr[cpu,hub]"
python
import onnx_asr

model = onnx_asr.load_model(
    "nemo-conformer-tdt",
    "calneymgp/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx-int8",
)
print(model.recognize("audio.wav", language="pt"))

Direct ONNX Runtime

The model follows the standard NVIDIA NeMo Conformer-TDT export layout. Any inference stack that supports nemo-conformer-tdt (e.g. NeMo's own nemo.collections.asr ONNX backend, custom Rust / C++ runtimes) should accept it after quantized op-set negotiation.


Reproducibility

python
from onnxruntime.quantization import quantize_dynamic, QuantType

# Encoder
quantize_dynamic(
    model_input="encoder-model.onnx",
    model_output="encoder-model.int8.onnx",
    weight_type=QuantType.QInt8,
    op_types_to_quantize=["MatMul"],
    use_external_data_format=False,
)

# Decoder + joint
quantize_dynamic(
    model_input="decoder_joint-model.onnx",
    model_output="decoder_joint-model.int8.onnx",
    weight_type=QuantType.QInt8,
    op_types_to_quantize=["MatMul"],
)

Environment: Python 3.11, onnxruntime 1.23.2, onnx 1.17.0, Windows 11.


Limitations & known caveats

  • —No formal WER benchmark on the quantized weights. Architecture is preserved and INT8 dynamic typically costs < 1 pp WER on Conformer ASR — but verify on your own pt-BR test set before production deployment.
  • —CPU-targeted. Dynamic INT8 may not accelerate GPU inference (CUDA EP). For GPU, prefer FP16 or static quantization with calibration.
  • —Single-file ONNX limit (~2 GB protobuf). Encoder fits comfortably at 839 MB; if you re-quantize differently and exceed 2 GB, switch use_external_data_format=True.
  • —Want smaller? INT4 weight-only via MatMul4BitsQuantizer reduces the encoder to ~280-380 MB but typically costs 1-3 pp WER and requires runtimes that support MatMulNBits op (ONNX Runtime ≥ 1.16). Whispering's transcribe-rs does not currently support INT4.

License

Inherits CC-BY-4.0 from the base model. Attribution is required for any downstream use.


Citation

If you use this quantized model, please cite the original work:

bibtex
@misc{parakeet-tdt-ptbr-tagarela,
  author = {Alefiury},
  title  = {Parakeet TDT 0.6B v3 pt-BR TAGARELA (ONNX)},
  year   = {2025},
  url    = {https://huggingface.co/alefiury/parakeet-tdt-0.6b-v3-ptBR-TAGARELA-onnx}
}

@article{xu2023tdt,
  title  = {Efficient Sequence Transduction by Jointly Predicting Tokens and Durations},
  author = {Xu, Hainan and Jia, Fei and Majumdar, Somshubra and others},
  year   = {2023},
  eprint = {2304.06795},
  archivePrefix = {arXiv}
}

Acknowledgements

  • —[Alefiury](https://huggingface.co/alefiury) — TAGARELA fine-tuning + ONNX export
  • —[Alexandre Acff](https://huggingface.co/alexandreacff) — pt-BR base fine-tune
  • —[NVIDIA NeMo team](https://github.com/NVIDIA/NeMo) — Parakeet architecture, FastConformer encoder, TDT decoder, multilingual pretraining
  • —[istupakov/onnx-asr](https://github.com/istupakov/onnx-asr) — reference ONNX inference framework
  • —[EpicenterHQ / Whispering](https://github.com/EpicenterHQ/epicenter) — local-first transcription app this quantization targets