datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
African_voices_yoruba
🇳🇬 WaZoBiaSpeech: 500+ Hour Yoruba (yor) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Yoruba (yor). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_yoruba.yoruba_audio_translatedThis is a copy of odunola/Yoruba_translate_preprocessed, the only difference is, it's already splitted into train & test. Awesome credits to her, her license applies too.
yoruba-speech-text-parallel
Yoruba Speech-Text Parallel Dataset
Dataset Description
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Yoruba - yo
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.s2tt-yoruba-english9jalingo-reviewed-yorubaafrispeech_yorubanaija-voices-yoruba-split_2-3naija-voices-yoruba-split_0-6naija-voices-yoruba-split_0-8naija-voices-yoruba-split_0-5final_yorubanaija-voices-yoruba-split_0-7yfacc_yorubanaija-voices-yoruba-split_2-7naija-voices-yoruba-split_2-0naija-voices-yoruba-split_2-2naija-voices-yoruba-split_0-4naija-voices-yoruba-split_1-0naija-voices-yoruba-split_2-4naija-voices-yoruba-split_2-6naija-voices-yoruba-split_1-2naija-voices-yoruba-split_1-6naija-voices-yoruba-split_2-8naija-voices-yoruba-split_1-4naija-voices-yoruba-split_0-9Additional_Yoruba_Datanaija-voices-yoruba-split_1-3yoruba-speech-transcribed
Yoruba Spontaneous Speech, Transcribed — Silencio
Spontaneous Yoruba with human-validated, fully tone-marked transcription and word-level alignment. 49 clips from 30 distinct speakers, 29 of them from Nigeria, across Ibadan, Lagos, Oyo and Nigerian Standard varieties. Transcripts keep the tonal diacritics and under-dots, and keep the Yoruba–English code-switching as it was spoken.
Hours
0.53
Clips
49
Speakers
30
Countries
2
Speaker origin regions
7
Native… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/yoruba-speech-transcribed.naija-voices-yoruba-split_1-7yoruba-dataset
