rmndrnts/wikipedia_asr_splitted
Synthetic dataset based on highly specialised texts from Wikipedia. This version is splitted by category. Voiced using Yandex SpeechKit with random voices, roles and speech rate. Can be used to evaluate ASR models not trained on given domains and to identify areas that the model does not handle well. Non-splitted version can be found here: https://huggingface.co/datasets/rmndrnts/wikipedia_asr
Synthetic dataset based on highly specialised texts from Wikipedia. This version is splitted by category. Voiced using Yandex SpeechKit with random voices, roles and speech rate. Can be used to evaluate ASR models not trained on given domains and to identify areas that the model does not handle well. Non-splitted version can be found here: https://huggingface.co/datasets/rmndrnts/wikipedia_asr
