CoolFace
Modelpublic

Yiivgeny/parakeet-tdt-0.6b-v3-sherpa-onnx-fp16

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes
Model Card

Parakeet TDT 0.6B v3 sherpa-onnx FP16

This repository contains a sherpa-onnx-compatible FP16 export of `nvidia/parakeet-tdt-0.6b-v3`.

It is intended for runtimes that expect the sherpa-onnx transducer layout:

text
tokens.txt
bpe.vocab
encoder.onnx
decoder.onnx
joiner.onnx

The encoder, decoder, and joiner weights are FP16. Public ONNX inputs and outputs are kept as FP32 via inserted casts, so sherpa-onnx can feed standard float32 features and decoder states.

Unlike the matching FP32 export, this FP16 package embeds the encoder weights inside encoder.onnx and does not require encoder.weights. An external-data FP16 encoder was loadable in Python ONNX Runtime, but the bundled sherpa-onnx websocket runtime used for OpenWhispr failed during startup with that layout. The embedded encoder passed both upstream ONNX testing and sherpa-onnx websocket smoke testing.

Downloads

The model files are available both directly in this repository and as a compressed archive:

text
sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-fp16.tar.bz2

Checksums are in:

text
SHA256SUMS

Verify after download:

bash
sha256sum -c SHA256SUMS

On macOS, use:

bash
shasum -a 256 -c SHA256SUMS

Usage With sherpa-onnx

Extract the archive:

bash
tar -xjf sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-fp16.tar.bz2
cd sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-fp16

Start the offline websocket server:

bash
sherpa-onnx-ws \
  --tokens=./tokens.txt \
  --encoder=./encoder.onnx \
  --decoder=./decoder.onnx \
  --joiner=./joiner.onnx \
  --port=6006 \
  --num-threads=4

Export

This export was produced from nvidia/parakeet-tdt-0.6b-v3 by first using the sherpa-onnx NeMo export script for Parakeet TDT 0.6B v3, then converting the exported ONNX models to FP16 with ONNX Runtime's float16 converter while preserving FP32 public IO types.

Expected FP16 output files:

text
encoder.onnx
decoder.onnx
joiner.onnx
tokens.txt
bpe.vocab

The export was validated with onnx.checker, upstream sherpa-onnx test_onnx.py, and the OpenWhispr bundled sherpa-onnx websocket server by transcribing the JFK sample en.wav.

Notes

This is a non-quantized FP16 export. It is smaller than the matching FP32 export but larger than int8 packages. A related public package, ako101/parakeet-tdt-0.6b-v3-sherpa-onnx-fp16, uses an FP16 encoder with int8 decoder and joiner; this package keeps all three transducer components in FP16.

License

The base model is released by NVIDIA under the CC BY 4.0 license. This converted ONNX export keeps the same license.