CoolFace
Modelpublic

rhasspy/qwen3-asr-0.6b-onnx-int4

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes56downloads
Model Card

Qwen3-ASR-0.6B ONNX — int4 encoder

`andrewleech/qwen3-asr-0.6b-onnx` with the audio encoder requantized to int4. The decoders, embedding table, config, and tokenizer are that repo's published int4 files, unchanged.

Ultimately derived from `Qwen/Qwen3-ASR-0.6B`. Apache-2.0 throughout.

Why

The upstream repo's encoder.int4.onnx is misleadingly named: it holds FP32 weights (746 MB). Quantizing it for real shrinks the package, and on ARM it cuts latency too.

upstreamthis
encoder746 MB (FP32)121 MB (int4)
total on disk2.0 GB1.4 GB
Raspberry Pi 5, 3.5 s utterance, 4 threads1.577 s1.36–1.49 s
Ryzen 9 5950X, 3.5 s utterance, 16 threads0.408 s0.411 s
Pi 5 peak RSS, short utterance—1.7 GB

The encoder is compute-bound, so the speedup is ARM-only; on x86 this is purely a memory saving. Transcripts matched the FP32 encoder on 9 of 10 Home Assistant test clips — the tenth was an entity name both get wrong without a context prompt and both get right with one.

Files

FileDescription
encoder.int4.onnx (+ .data)Audio encoder, int4 MatMulNBits
decoder_init.int4.onnxDecoder prefill; takes input_ids, emits logits + KV cache
decoder_step.int4.onnxAutoregressive step; takes input_embeds + KV cache
decoder_weights.int4.dataShared external weights for both decoders
embed_tokens.binToken embeddings [151936, 1024], float16
config.json, tokenizer.jsonArchitecture config, special tokens, mel params, tokenizer

Keep embed_tokens.bin in fp16 and cast one row per generated token; casting the whole table at load costs ~300 MB of RSS for nothing.

Inference

  1. 1.Log-mel spectrogram (Whisper parameters: 16 kHz, 128 bins, n_fft 400, hop 160, Hann, Slaney mel, 0–8 kHz)
  2. 2.encoder.int4.onnx → audio features
  3. 3.Build prompt ids, then decoder_init.int4.onnx with input_ids, position_ids, audio_features, audio_offset
  4. 4.Greedy loop on decoder_step.int4.onnx until <|im_end|> or <|endoftext|>
  5. 5.Drop everything up to and including <asr_text> (the language preamble)

The prompt is the Qwen chat template:

<|im_start|>system\n{context}<|im_end|>\n
<|im_start|>user\n<|audio_start|>{audio_pad × N}<|audio_end|><|im_end|>\n
<|im_start|>assistant\n{language {Name}<asr_text>}

Note the token ids for the words system and user are 8948 and 872. The upstream reference src/prompt.py hardcodes 9125 and 882, which decode to " Current" and " time".

Context biasing

Free-form text in the system turn biases decoding toward specific spellings — useful for smart-home entity names:

system: Vocabulary: Ecobee, office lamp.

"What's the temperature of the incubator?" → "What's the temperature of the Ecobee?"

Verified clean with a 77-token entity list: no dropped outputs, no prompt echo. (The int8 export at OpenVoiceOS/qwen3-asr-0.6b-onnx degenerates under the same load — empty output, or echoing the vocabulary list back as the transcript. int4 MatMulNBits' per-group scales handle the decoder's outlier weights that per-tensor int8 does not.)

Reproducing the encoder

python
import onnx
from onnxruntime.quantization.matmul_nbits_quantizer import (
    MatMulNBitsQuantizer, RTNWeightOnlyQuantConfig)
from onnxruntime.quantization.quant_utils import QuantFormat

q = MatMulNBitsQuantizer(
    model=onnx.load("encoder.int4.onnx"),   # the FP32-weighted file from upstream
    block_size=64, is_symmetric=False, accuracy_level=4,
    algo_config=RTNWeightOnlyQuantConfig(quant_format=QuantFormat.QOperator),
)
q.process()
q.model.save_model_to_file("encoder.int4.onnx", use_external_data_format=True)

block_size=64 / accuracy_level=4 match the recipe the decoders were built with. Don't change them casually — upstream measured block_size=32 at the same accuracy level producing 99.98% WER, and requantizing the decoder with these settings produced empty output on both x86 and ARM.