datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yoruba-speech-text-parallel
Yoruba Speech-Text Parallel Dataset
Dataset Description
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Yoruba - yo
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.yoruba-speech-data
Yoruba Speech Data (Pooled)
A ~1,039-hour pooled Yoruba speech corpus, combining four independent sources into
one consistently-formatted dataset for speech modeling (TTS / ASR). Built to give a
Yoruba TTS finetune enough scale and diversity to move past what a single ~100-hour
source can teach a model — more speakers, more domains, more of the language.
Sources
Source
Clips
Hours
Style
DSN African Voices (Data Science Nigeria)
144,628
309.0 h… See the full description on the dataset page: https://huggingface.co/datasets/Professor/yoruba-speech-data.yoruba-speech-transcribed
Yoruba Spontaneous Speech, Transcribed — Silencio
Spontaneous Yoruba with human-validated, fully tone-marked transcription and word-level alignment. 49 clips from 30 distinct speakers, 29 of them from Nigeria, across Ibadan, Lagos, Oyo and Nigerian Standard varieties. Transcripts keep the tonal diacritics and under-dots, and keep the Yoruba–English code-switching as it was spoken.
Hours
0.53
Clips
49
Speakers
30
Countries
2
Speaker origin regions
7
Native… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/yoruba-speech-transcribed.yoruba-second-sbpn-demucs-20260826
yoruba-second-sbpn-demucs-20260826
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-second-sbpn-demucs-20260826.yoruba-datasold-sbpn-demucs-20260825
yoruba-datasold-sbpn-demucs-20260825
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-datasold-sbpn-demucs-20260825.farahan-yoruba-elder-corpus
Apakose Ezekiel Imoleayo — Farahàn
Appear. Be found.
Lagos, Nigeria · UNILAG Yoruba Studies ·
Graduating 2030
What I Build
Yoruba Oral Knowledge Corpus
I document primary-source Yoruba knowledge
directly from elder speakers in Lagos and
southwest Nigeria. Structured interviews
covering proverbs (Owe), oral history (Itan),
praise poetry (Oriki), and cultural knowledge
systems that no web scrape produces.
What makes this corpus different:… See the full description on the dataset page: https://huggingface.co/datasets/Apakose-Ezekiel/farahan-yoruba-elder-corpus.yoruba-bolanle-sbpn-demucs-20260825
yoruba-bolanle-sbpn-demucs-20260825
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-bolanle-sbpn-demucs-20260825.9javoice-yoruba
9jaVoice Consent-1
153 clips. 29.4 minutes, which is 0.5 hours. 13 speakers. Yoruba. FLAC, 48 kHz, mono, 16-bit. CC BY-NC 4.0.
Read-aloud speech, recorded by paid contributors on their own phones. Use it for evaluation, for fine-tuning, and as a reference set when you want to find out whether a model handles Nigerian speech at all. It is too small to pretrain on and we are not going to pretend otherwise.
Every clip here carries its own consent record. The contributor ticked an… See the full description on the dataset page: https://huggingface.co/datasets/9jatesters/9javoice-yoruba.
