Yiivgeny/parakeet-tdt-0.6b-v3-sherpa-onnx-fp16
Parakeet TDT 0.6B v3 sherpa-onnx FP16
This repository contains a sherpa-onnx-compatible FP16 export of `nvidia/parakeet-tdt-0.6b-v3`.
It is intended for runtimes that expect the sherpa-onnx transducer layout:
tokens.txt
bpe.vocab
encoder.onnx
decoder.onnx
joiner.onnxThe encoder, decoder, and joiner weights are FP16. Public ONNX inputs and outputs are kept as FP32 via inserted casts, so sherpa-onnx can feed standard float32 features and decoder states.
Unlike the matching FP32 export, this FP16 package embeds the encoder weights inside encoder.onnx and does not require encoder.weights. An external-data FP16 encoder was loadable in Python ONNX Runtime, but the bundled sherpa-onnx websocket runtime used for OpenWhispr failed during startup with that layout. The embedded encoder passed both upstream ONNX testing and sherpa-onnx websocket smoke testing.
Downloads
The model files are available both directly in this repository and as a compressed archive:
sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-fp16.tar.bz2Checksums are in:
SHA256SUMSVerify after download:
sha256sum -c SHA256SUMSOn macOS, use:
shasum -a 256 -c SHA256SUMSUsage With sherpa-onnx
Extract the archive:
tar -xjf sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-fp16.tar.bz2
cd sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-fp16Start the offline websocket server:
sherpa-onnx-ws \
--tokens=./tokens.txt \
--encoder=./encoder.onnx \
--decoder=./decoder.onnx \
--joiner=./joiner.onnx \
--port=6006 \
--num-threads=4Export
This export was produced from nvidia/parakeet-tdt-0.6b-v3 by first using the sherpa-onnx NeMo export script for Parakeet TDT 0.6B v3, then converting the exported ONNX models to FP16 with ONNX Runtime's float16 converter while preserving FP32 public IO types.
Expected FP16 output files:
encoder.onnx
decoder.onnx
joiner.onnx
tokens.txt
bpe.vocabThe export was validated with onnx.checker, upstream sherpa-onnx test_onnx.py, and the OpenWhispr bundled sherpa-onnx websocket server by transcribing the JFK sample en.wav.
Notes
This is a non-quantized FP16 export. It is smaller than the matching FP32 export but larger than int8 packages. A related public package, ako101/parakeet-tdt-0.6b-v3-sherpa-onnx-fp16, uses an FP16 encoder with int8 decoder and joiner; this package keeps all three transducer components in FP16.
License
The base model is released by NVIDIA under the CC BY 4.0 license. This converted ONNX export keeps the same license.
