rmndrnts/wikipedia_asr_splitted
Synthetic dataset based on highly specialised texts from Wikipedia. This version is splitted by category. Voiced using Yandex SpeechKit with random voices, roles and speech rate. Can be used to evaluate ASR models not trained on given domains and to identify areas that the model does not handle well. Non-splitted version can be found here: https://huggingface.co/datasets/rmndrnts/wikipedia_asr
0363
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face