rhasspy/qwen3-asr-0.6b-onnx-int4
Qwen3-ASR-0.6B ONNX — int4 encoder
`andrewleech/qwen3-asr-0.6b-onnx` with the audio encoder requantized to int4. The decoders, embedding table, config, and tokenizer are that repo's published int4 files, unchanged.
Ultimately derived from `Qwen/Qwen3-ASR-0.6B`. Apache-2.0 throughout.
Why
The upstream repo's encoder.int4.onnx is misleadingly named: it holds FP32 weights (746 MB). Quantizing it for real shrinks the package, and on ARM it cuts latency too.
The encoder is compute-bound, so the speedup is ARM-only; on x86 this is purely a memory saving. Transcripts matched the FP32 encoder on 9 of 10 Home Assistant test clips — the tenth was an entity name both get wrong without a context prompt and both get right with one.
Files
Keep embed_tokens.bin in fp16 and cast one row per generated token; casting the whole table at load costs ~300 MB of RSS for nothing.
Inference
- Log-mel spectrogram (Whisper parameters: 16 kHz, 128 bins, n_fft 400, hop 160, Hann, Slaney mel, 0–8 kHz)
encoder.int4.onnx→ audio features- Build prompt ids, then
decoder_init.int4.onnxwithinput_ids,position_ids,audio_features,audio_offset - Greedy loop on
decoder_step.int4.onnxuntil<|im_end|>or<|endoftext|> - Drop everything up to and including
<asr_text>(the language preamble)
The prompt is the Qwen chat template:
<|im_start|>system\n{context}<|im_end|>\n
<|im_start|>user\n<|audio_start|>{audio_pad × N}<|audio_end|><|im_end|>\n
<|im_start|>assistant\n{language {Name}<asr_text>}Note the token ids for the words system and user are 8948 and 872. The upstream reference src/prompt.py hardcodes 9125 and 882, which decode to " Current" and " time".
Context biasing
Free-form text in the system turn biases decoding toward specific spellings — useful for smart-home entity names:
system: Vocabulary: Ecobee, office lamp."What's the temperature of the incubator?" → "What's the temperature of the Ecobee?"
Verified clean with a 77-token entity list: no dropped outputs, no prompt echo. (The int8 export at OpenVoiceOS/qwen3-asr-0.6b-onnx degenerates under the same load — empty output, or echoing the vocabulary list back as the transcript. int4 MatMulNBits' per-group scales handle the decoder's outlier weights that per-tensor int8 does not.)
Reproducing the encoder
import onnx
from onnxruntime.quantization.matmul_nbits_quantizer import (
MatMulNBitsQuantizer, RTNWeightOnlyQuantConfig)
from onnxruntime.quantization.quant_utils import QuantFormat
q = MatMulNBitsQuantizer(
model=onnx.load("encoder.int4.onnx"), # the FP32-weighted file from upstream
block_size=64, is_symmetric=False, accuracy_level=4,
algo_config=RTNWeightOnlyQuantConfig(quant_format=QuantFormat.QOperator),
)
q.process()
q.model.save_model_to_file("encoder.int4.onnx", use_external_data_format=True)block_size=64 / accuracy_level=4 match the recipe the decoders were built with. Don't change them casually — upstream measured block_size=32 at the same accuracy level producing 99.98% WER, and requantizing the decoder with these settings produced empty output on both x86 and ARM.
