CoolFace
Modelpublic

shigedonsan/qwen3-asr-1.7b-sherpa-onnx-int8-4096

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes
Model Card

Qwen3-ASR 1.7B for sherpa-onnx — INT8 / 4096 context

An offline ONNX conversion of Qwen3-ASR-1.7B for CPU use with sherpa-onnx. This is an independent community conversion, not an official Qwen release.

  • —Encoder and convolution frontend: FP16
  • —Decoder: mixed INT8/FP32
  • —Exported max_total_len: 4096

Files

  • —conv_frontend.onnx
  • —encoder.onnx and encoder.onnx.data
  • —decoder.int8.onnx and decoder.int8.onnx.data
  • —tokenizer/

Usage

python
import sherpa_onnx

recognizer = sherpa_onnx.OfflineRecognizer.from_qwen3_asr(
    conv_frontend="conv_frontend.onnx",
    encoder="encoder.onnx",
    decoder="decoder.int8.onnx",
    tokenizer="tokenizer",
    provider="cpu",
    num_threads=4,
    max_total_len=4096,
    max_new_tokens=1024,
)
stream = recognizer.create_stream()
stream.accept_waveform(16000, audio_float32_mono)
recognizer.decode_stream(stream)
print(stream.result.text)

Optional hotwords: stream.set_option("hotwords", "Kubernetes,OpenShift").

Notes

  • —Validated with sherpa-onnx 1.13.8; max_new_tokens=1024 was used.
  • —CPU performance depends on hardware and audio length.
  • —Detailed provenance is in manifest.json and NOTICE.

License

Apache-2.0. Based on Qwen3-ASR; see LICENSE and NOTICE for attribution. The conversion is community-maintained and is not an official Qwen release.