a1700878328/qwen3-asr-1.7b-onnx-fp16
Qwen3-ASR-1.7B ONNX FP16
GPU-oriented ONNX Runtime export of Qwen3-ASR-1.7B.
This is a split, autoregressive ONNX model, not a sherpa-onnx package:
encoder.onnx: native FP16 audio encoderdecoder_init.onnx: prompt prefill and initial KV cachedecoder_step.onnx: cached one-token decodingdecoder_weights.data: shared external decoder weightsembed_tokens.bin: FP16 token embeddings
The decoder weights, matrix multiplications, embeddings, logits, and KV cache use FP16. Rotary embedding and RMSNorm subgraphs remain FP32 because the upstream Qwen implementation intentionally performs these numerically sensitive operations in FP32. This mixed-compute detail prevents the degenerate output produced by blindly converting every node to FP16.
Requirements
- NVIDIA GPU with FP16 support
- CUDA 12.x and cuDNN 9.x
- ONNX Runtime GPU 1.23.2 or newer
- About 9 GiB of free VRAM with three independent ORT sessions loaded
Use both execution providers, but fail startup if CUDA is not first. ONNX Runtime can otherwise silently continue on CPU when CUDA libraries are missing:
import onnxruntime as ort
ort.preload_dlls()
providers = ["CUDAExecutionProvider", "CPUExecutionProvider"]
session = ort.InferenceSession("encoder.onnx", providers=providers)
if session.get_providers()[0] != "CUDAExecutionProvider":
raise RuntimeError("CUDAExecutionProvider failed to initialize")CPUExecutionProvider is retained for small shape and indexing operations. Profiling confirmed that all MatMul and Conv nodes execute with CUDA.
Measured results
Measured locally on an RTX 5070 Ti with greedy decoding:
On the two long Japanese samples, normalized character error rate was 3.71% for this FP16 export and 3.62% for the FP32/PyTorch reference. Their decoded texts differed by one normalized character across 1,026 generated characters. These are small local samples, not a general benchmark.
The three loaded ORT sessions consumed about 8.8 GiB of additional VRAM in this test environment. Runtime memory varies with CUDA/ORT versions and audio length.
File interface
The model uses FP16 interfaces for decoder audio features, token embeddings, logits, and KV caches. A consumer written for an FP32 ONNX export must inspect the declared input types or cast these tensors to FP16. config.json records the embedding, decoder, and mixed-compute precision.
The decoder files intentionally share decoder_weights.data. Keep these three files in the same directory:
decoder_init.onnx
decoder_step.onnx
decoder_weights.dataLimitations
- Greedy offline decoding only in the included graph interface
- No timestamps or forced alignment
- No dynamic batching interface
- Long-form chunking and multi-user request scheduling belong in the server
- GPU throughput can be improved further with I/O binding and server-side batching
Attribution
The original weights and tokenizer are from Qwen/Qwen3-ASR-1.7B, licensed under Apache-2.0. This repository contains a converted ONNX representation.
