datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/claudefitz/WaxalNLP.LOD_Claude
LOD_Claude Dataset
Dataset Description
LOD_Claude is a Luxembourgish speech dataset containing audio recordings paired with transcriptions. The audio features a synthetic voice named Claude reading example sentences from the LOD (Lëtzebuerger Online Dictionnaire) available at lod.lu.
Dataset Statistics
Total samples: 39,034
Training samples: 37,084
Validation samples: 1,950
Language: Luxembourgish (Lëtzebuergesch)
Audio format: WAV files
Sample rate: 24,000… See the full description on the dataset page: https://huggingface.co/datasets/ZLSCompLing/LOD_Claude.
