yuriyvnv/Qwen3-ASR-1.7B-NL
ποΈ Qwen3-ASR-1.7B-NL β Dutch Speech Recognition
<div align="center"> <img src="https://img.shields.io/badge/Parameters-1.7B-red" alt="1.7B Parameters"> <img src="https://img.shields.io/badge/Modality-Speech%20%E2%86%92%20Text-purple" alt="Speech to Text"> <img src="https://img.shields.io/badge/Language-Dutch-green" alt="Dutch"> <img src="https://img.shields.io/badge/Task-ASR-blue" alt="Automatic Speech Recognition"> <img src="https://img.shields.io/badge/Base-Qwen3--ASR--1.7B-orange" alt="Base model"> <img src="https://img.shields.io/badge/Precision-bf16-lightgrey" alt="bf16"> <img src="https://img.shields.io/badge/License-Apache--2.0-yellow" alt="Apache-2.0"> </div>
<br/>
A Dutch-specialised automatic speech recognition (ASR) model, fine-tuned from Qwen/Qwen3-ASR-1.7B. It outputs cased, punctuated Dutch text and works as a drop-in replacement for the base model.
On Common Voice 22 (test) it reaches 5.28% WER (down from 6.68% zero-shot, -21.0% relative).
π Results
WER and CER on held-out Common Voice test sets β same samples, same protocol, no test-time tricks. "Zero-shot" is the base Qwen/Qwen3-ASR-1.7B called with language="Dutch". The fine-tuned numbers are bold.
Lower is better. Both held-out test sets see roughly a one-third relative reduction in word error rate versus the already-strong base model.
π¬ Reproducibility note. Both the zero-shot baseline and the fine-tuned numbers above were measured with the same evaluation function (train_qwen3_asr.evaluate_model), the same greedy decoding settings, and the same reference normalisation (see next section). This is an apples-to-apples comparison.π§Ή Reference / target normalisation
Common Voice transcripts are crowd-sourced and inconsistent in casing and trailing punctuation. To give the model a clean, predictable target distribution we apply a small, deterministic written-form normalisation to every reference at load time, both during training and during evaluation:
- Capitalise the first letter if it is lowercase.
- Collapse trailing dots β any sequence of
.,β¦,..,...at the end is replaced with a single.. - Append a terminal period if the sentence does not already end in terminal punctuation (
. ! ? β¦) or a closing bracket / quote () ] } " 'etc.).
The exact function lives in src/evaluation/score_written_form.py of the project repository. Concretely:
Because the same normalisation is applied to references used for the zero-shot baseline above, the gain reported in the results table reflects the fine-tune itself β not a metric quirk caused by mismatched references.
π How to use
Install the official qwen-asr package, then load this model exactly the same way you would load the base Qwen3-ASR:
pip install qwen-asrimport torch
from qwen_asr import Qwen3ASRModel
model = Qwen3ASRModel.from_pretrained(
"yuriyvnv/Qwen3-ASR-1.7B-NL",
dtype=torch.bfloat16,
device_map="cuda:0",
)
result = model.transcribe(audio="audio.wav", language="Dutch")
print(result[0].text)Batch inference, automatic language detection, streaming, and vLLM serving all work identically to the base model β see the upstream Qwen3-ASR documentation for details.
π οΈ Training
Dataset: yuriyvnv/synthetic_transcript_nl + fsicoli/common_voice_22_0 (nl) β the full synthetic OpenAI-TTS Dutch corpus (~34.9k clips) concatenated with the Common Voice 22 Dutch train split, shuffled with the run seed. After duration filtering and transcript-length filtering: 45,898 training samples and 11,000 validation samples.
Recipe: follows the official QwenLM SFT recipe with our local hyperparameters:
Trained on a single H100. The best checkpoint was selected by validation loss.
β οΈ Limitations
- Trained on Common Voice β read-speech dominated. Conversational, overlapping-speaker, far-field, or strongly accented audio may degrade accuracy.
- Outputs Dutch text. Cross-lingual or code-switched audio is not targeted.
- Punctuation and casing are best-effort and inherit the inconsistencies of the Common Voice reference transcripts (mitigated, but not eliminated, by the normalisation step above).
π Acknowledgements
This model would not exist without the work of others. Thank you to:
- The Qwen team at Alibaba Cloud for releasing Qwen3-ASR-1.7B β the backbone of this fine-tune β together with a clean, reproducible SFT recipe and the Qwen3-ASR Technical Report.
- The Mozilla Common Voice community for collecting and releasing the Dutch speech corpus used for training and evaluation (Common Voice 22, Common Voice 17 mirror).
- Every contributor who recorded, validated, or transcribed a clip in Common Voice. This model is, very literally, your voices.
π Citation
If this model is useful in your work, please cite the base Qwen3-ASR report:
@article{qwen3asr2025,
title = {Qwen3-ASR Technical Report},
author = {Qwen Team},
year = {2025},
url = {https://arxiv.org/abs/2601.21337}
}