CoolFace
Datasetpublic

rmndrnts/wikipedia_asr_splitted

Synthetic dataset based on highly specialised texts from Wikipedia. This version is splitted by category. Voiced using Yandex SpeechKit with random voices, roles and speech rate. Can be used to evaluate ASR models not trained on given domains and to identify areas that the model does not handle well. Non-splitted version can be found here: https://huggingface.co/datasets/rmndrnts/wikipedia_asr

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes341downloads
Dataset Card

Synthetic dataset based on highly specialised texts from Wikipedia. This version is splitted by category. Voiced using Yandex SpeechKit with random voices, roles and speech rate. Can be used to evaluate ASR models not trained on given domains and to identify areas that the model does not handle well. Non-splitted version can be found here: https://huggingface.co/datasets/rmndrnts/wikipedia_asr