datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
African_voices_yoruba
🇳🇬 WaZoBiaSpeech: 500+ Hour Yoruba (yor) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Yoruba (yor). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_yoruba.yoruba_dataset_encodedyoruba_audio_translatedThis is a copy of odunola/Yoruba_translate_preprocessed, the only difference is, it's already splitted into train & test. Awesome credits to her, her license applies too.
yoruba-speech-text-parallel
Yoruba Speech-Text Parallel Dataset
Dataset Description
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Yoruba - yo
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.9jalingo-reviewed-yorubas2tt-yoruba-englishyoruba-cfm-latentsyoruba_dataset_encoded_repo20to30afrispeech_yorubafinal_yorubayoruba-normalization-pairs
Normalization pairs dataset
What this is
24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code.
The library
This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.naija-voices-yoruba-split_2-3yoruba
Yoruba Dataset
The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Yoruba words.
Data Types
Synthetic sentences (400k): Basic vocabulary combinations
Pattern variations (200k): Template-based grammatical structures
Conversations (150k): Interactive dialogue examples
Q&A pairs (100k): Knowledge and reasoning tasks
Translations (80k): Yoruba-English bidirectional pairs
Grammar examples (40k): Verb… See the full description on the dataset page: https://huggingface.co/datasets/0xnu/yoruba.yfacc_yorubayoruba-setnaija-voices-yoruba-split_0-8naija-voices-yoruba-split_0-6naija-voices-yoruba-split_0-7naija-voices-yoruba-split_2-7naija-voices-yoruba-split_0-5naija-voices-yoruba-split_2-0naija-voices-yoruba-split_1-0naija-voices-yoruba-split_2-2naija-voices-yoruba-split_2-6naija-voices-yoruba-split_0-4naija-voices-yoruba-split_1-2naija-voices-yoruba-split_1-6naija-voices-yoruba-split_2-4Additional_Yoruba_Dataenglish-yoruba_sentence-pairs_mt560
English-Yoruba Parallel Dataset
This dataset contains parallel sentences in English and Yoruba (Nigeria).
Dataset Information
Language Pair: English ↔ Yoruba
Language Code: yor
Country: Nigeria
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-yoruba_sentence-pairs_mt560.
