CoolFace
Modelpublic

a1700878328/qwen3-asr-1.7b-onnx-fp16

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes9downloads
Model Card

Qwen3-ASR-1.7B ONNX FP16

GPU-oriented ONNX Runtime export of Qwen3-ASR-1.7B.

This is a split, autoregressive ONNX model, not a sherpa-onnx package:

  • —encoder.onnx: native FP16 audio encoder
  • —decoder_init.onnx: prompt prefill and initial KV cache
  • —decoder_step.onnx: cached one-token decoding
  • —decoder_weights.data: shared external decoder weights
  • —embed_tokens.bin: FP16 token embeddings

The decoder weights, matrix multiplications, embeddings, logits, and KV cache use FP16. Rotary embedding and RMSNorm subgraphs remain FP32 because the upstream Qwen implementation intentionally performs these numerically sensitive operations in FP32. This mixed-compute detail prevents the degenerate output produced by blindly converting every node to FP16.

Requirements

  • —NVIDIA GPU with FP16 support
  • —CUDA 12.x and cuDNN 9.x
  • —ONNX Runtime GPU 1.23.2 or newer
  • —About 9 GiB of free VRAM with three independent ORT sessions loaded

Use both execution providers, but fail startup if CUDA is not first. ONNX Runtime can otherwise silently continue on CPU when CUDA libraries are missing:

python
import onnxruntime as ort

ort.preload_dlls()
providers = ["CUDAExecutionProvider", "CPUExecutionProvider"]
session = ort.InferenceSession("encoder.onnx", providers=providers)
if session.get_providers()[0] != "CUDAExecutionProvider":
    raise RuntimeError("CUDAExecutionProvider failed to initialize")

CPUExecutionProvider is retained for small shape and indexing operations. Profiling confirmed that all MatMul and Conv nodes execute with CUDA.

Measured results

Measured locally on an RTX 5070 Ti with greedy decoding:

AudioDurationEncoderDecodeTotal RTF
Japanese short sample15.1 s0.73 s3.96 s0.31
Japanese narration108.9 s0.89 s40.99 s0.38
Japanese dialogue104.7 s0.05 s warm40.83 s0.39

On the two long Japanese samples, normalized character error rate was 3.71% for this FP16 export and 3.62% for the FP32/PyTorch reference. Their decoded texts differed by one normalized character across 1,026 generated characters. These are small local samples, not a general benchmark.

The three loaded ORT sessions consumed about 8.8 GiB of additional VRAM in this test environment. Runtime memory varies with CUDA/ORT versions and audio length.

File interface

The model uses FP16 interfaces for decoder audio features, token embeddings, logits, and KV caches. A consumer written for an FP32 ONNX export must inspect the declared input types or cast these tensors to FP16. config.json records the embedding, decoder, and mixed-compute precision.

The decoder files intentionally share decoder_weights.data. Keep these three files in the same directory:

text
decoder_init.onnx
decoder_step.onnx
decoder_weights.data

Limitations

  • —Greedy offline decoding only in the included graph interface
  • —No timestamps or forced alignment
  • —No dynamic batching interface
  • —Long-form chunking and multi-user request scheduling belong in the server
  • —GPU throughput can be improved further with I/O binding and server-side batching

Attribution

The original weights and tokenizer are from Qwen/Qwen3-ASR-1.7B, licensed under Apache-2.0. This repository contains a converted ONNX representation.