CoolFace
Modelpublic

mults/mova-asr-uk-citrinet512-int8

sourceHugging Facebsd-3-clauseupdated 2mo agoView on Hugging Face
0likes
Model Card

Ukrainian Citrinet-512 — sherpa-onnx ONNX int8 (on-device / mobile)

What it is: neongeckocom/stt_uk_citrinet_512_gamma_0_25 (a NeMo Citrinet-512 Ukrainian speech-recognition model by Neon AI) exported to ONNX, dynamically quantized to int8 (36 MB), with the metadata sherpa-onnx expects for its nemo_ctc offline recognizer class.

Goal: fully offline Ukrainian speech-to-text on phones and other on-device targets — small, fast (non-autoregressive CTC decode), no NeMo/PyTorch needed at inference time. Verified on Android via sherpa-onnx.

Usage (sherpa-onnx)

python
import sherpa_onnx
recognizer = sherpa_onnx.OfflineRecognizer.from_nemo_ctc(
    model="model.int8.onnx", tokens="tokens.txt", num_threads=4)
stream = recognizer.create_stream()
stream.accept_waveform(16000, samples_float32)  # 16 kHz mono, [-1, 1]
recognizer.decode_stream(stream)
print(stream.result.text)

Export recipe: NeMo model.export() → sherpa-onnx metadata (normalize_type=per_feature, subsampling_factor=8, model_type=EncDecCTCModelBPE) → onnxruntime.quantization.quantize_dynamic (QUInt8) → tokens.txt from the decoder vocabulary + trailing <blk>, per the k2-fsa NeMo export guide.

Credits & license

  • —Model weights: Neon AI — stt_uk_citrinet_512_gamma_0_25 (BSD-3-Clause), trained on openstt-uk; based on NVIDIA's Citrinet architecture and English checkpoint (CC-BY-4.0).
  • —Conversion/quantization/packaging: the Mova project (offline travel translator), 2026.
  • —This export is redistributed under BSD-3-Clause, same as the source model.