andrewbawitlung/qwen3-asr-0.6b-mizonal3-E5-lus-v2026.06
07
Disclaimer / Notice: Details for these are in Peer Review and publications of the paper will be made available soon for more details.
<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->
qwen3-asr-0.6b-mizonal3-E5-lus-v2026.06
This model is a fine-tuned version of Qwen/Qwen3-ASR-0.6B on the MiZonal v3.0 dataset. Note: ~1 hour of conversational speech was added to this dataset version.
It achieves the following results on the evaluation set:
- Wer: 18.6414
- Cer: 4.2134
- Real Time Factor: 0.0718
Quick Inference
import torch
import librosa
from transformers import AutoProcessor, Qwen2AudioForConditionalGeneration
device = "cuda" if torch.cuda.is_available() else "cpu"
processor = AutoProcessor.from_pretrained("andrewbawitlung/qwen3-asr-0.6b-mizonal3-E5-lus-v2026.06")
model = Qwen2AudioForConditionalGeneration.from_pretrained("andrewbawitlung/qwen3-asr-0.6b-mizonal3-E5-lus-v2026.06").to(device)
audio, sr = librosa.load("your_audio.wav", sr=16000)
conversation = [
{"role": "user", "content": [
{"type": "audio", "audio_url": "your_audio.wav"},
{"type": "text", "text": "Transcribe the audio:"}
]}
]
text = processor.apply_chat_template(conversation, add_generation_prompt=True, tokenize=False)
inputs = processor(text=text, audios=[audio], return_tensors="pt", padding=True)
inputs.input_ids = inputs.input_ids.to(device)
with torch.no_grad():
generate_ids = model.generate(**inputs, max_length=256)
generate_ids = generate_ids[:, inputs.input_ids.size(1):]
transcription = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]
print(transcription)Model description
Experiment Configurations
This repository is part of a series of experiments. The different configurations are:
- E1 (Baseline): Standard training configuration.
- E2 (Noise): Training with background noise augmentation.
- E3 (Speed): Training with speed perturbation augmentation.
- E4 (SpecAug): Training with SpecAugment (time and frequency masking).
- E5 (Combined): Training with a combination of all augmentations.
All Models in this Family
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 2e-05
- trainbatchsize: 16
- evalbatchsize: 8
- seed: 42
- optimizer: OptimizerNames.ADAMWTORCHFUSED
- lrschedulertype: SchedulerType.LINEAR
- num_epochs: 8
