datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
az-asr-calls-190h
Azerbaijani call-centre speech, Soniox pseudo-labels
Real Azerbaijani call-centre audio transcribed by Soniox
stt-async-v5. 169,273 clips, 190.63 h, 8 kHz telephony.
The labels are machine output and carry a measured ceiling. Soniox scores
46.65% WER against human transcripts of this same kind of audio. A model
trained on these labels learns to agree with Soniox -- including where Soniox is
wrong, and including its habit of dropping words.
That is not hypothetical. Scoring the… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-calls-190h.az-asr-calls-human-140h
Azerbaijani call-centre speech, human-transcribed
Real Azerbaijani call-centre audio with transcripts written by people
listening to it. 29,982 clips, 139.75 h, 8 kHz telephony.
This is the most valuable corpus in the collection and the smallest. It is the
only one that is both human-transcribed and spontaneous telephony -- the register
Chinar-F8 actually targets. Adding 134.7 h of it to the training mix moved WER on
human-transcribed calls from 43.39% to 35.08%, the largest… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-calls-human-140h.
