datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
everyayah
EveryAyah Quran Dataset
Dataset Description
This dataset contains verse-by-verse (ayah) Quranic recitations from 4 professional reciters, with fully diacritized Arabic text aligned to each audio segment. The dataset is designed for automatic speech recognition (ASR), speaker identification, and other speech processing tasks focused on Classical Arabic and Quranic recitation.
Dataset Summary
Total Audio Files: 24,944 ayahs
Number of Reciters: 4
Audio Format:… See the full description on the dataset page: https://huggingface.co/datasets/Rdyh/everyayah.mediTalk-mm-rdy
Burmese Medical Speech Corpus for ASR and TTS
This dataset is under Non-Commercial (CC BY-NC-SA 4.0) License.
No commercial use is allowed and only for Research and Development.
We are still developing the corpus and for more information please contact via thuraaung(dot)ai(dot)mdy@gmail(dot)com.
More Information needed
