datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
procedural-engine-sounds
Procedural Engine Sounds Dataset (Official)
⚠️ Canonical Source
This is the original and official release of the Procedural Engine Sounds Dataset, created by Robin Doerfler:
Project page (start here): https://rdoerfler.github.io/procedural-engine-sounds-page/
Hugging Face repository: https://huggingface.co/datasets/rdoerfler/procedural-engine-sounds
Zenodo DOI (primary citation): https://doi.org/10.5281/zenodo.16883336
Other versions of this dataset on… See the full description on the dataset page: https://huggingface.co/datasets/rdoerfler/procedural-engine-sounds.rd3-da33aeeveryayah
EveryAyah Quran Dataset
Dataset Description
This dataset contains verse-by-verse (ayah) Quranic recitations from 4 professional reciters, with fully diacritized Arabic text aligned to each audio segment. The dataset is designed for automatic speech recognition (ASR), speaker identification, and other speech processing tasks focused on Classical Arabic and Quranic recitation.
Dataset Summary
Total Audio Files: 24,944 ayahs
Number of Reciters: 4
Audio Format:… See the full description on the dataset page: https://huggingface.co/datasets/Rdyh/everyayah.RDSThis dataset contains audio <> text pairs in bangru dialect of Haryanvi. filepath and text mapping are stored in metadata.csv file
rd316-mixed-sttTreinamentoRVCRdzinmediTalk-mm-rdy
Burmese Medical Speech Corpus for ASR and TTS
This dataset is under Non-Commercial (CC BY-NC-SA 4.0) License.
No commercial use is allowed and only for Research and Development.
We are still developing the corpus and for more information please contact via thuraaung(dot)ai(dot)mdy@gmail(dot)com.
More Information needed
tts_datamediTalk-mm-rdy-augrde-neste-authmediTalk-mm-rdy-classifyrdg5_indimediTalk-mm-rdy-st
Burmese Medical Speech Corpus for ASR and TTS
This dataset is under Non-Commercial (CC BY-NC-SA 4.0) License.
No commercial use is allowed and only for Research and Development.
We are still developing the corpus and for more information please contact via thuraaung(dot)ai(dot)mdy@gmail(dot)com.
More Information needed
mediTalk-mm-rdy-st-augmediTalk-mm-ae-rdy
Burmese Medical Speech Corpus for ASR and TTS
This dataset is under Non-Commercial (CC BY-NC-SA 4.0) License.
No commercial use is allowed and only for Research and Development.
We are still developing the corpus and for more information please contact via thuraaung(dot)ai(dot)mdy@gmail(dot)com.
mediTalk-mm-ae-rdy-augrdpd
