CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Africanvoice /African_voices_yoruba 🇳🇬 WaZoBiaSpeech: 500+ Hour Yoruba (yor) Corpus Version: 30 Nov 2025 NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release. 🌍 Dataset Overview WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Yoruba (yor). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_yoruba.audio100K<n<1M0 likes1.1k downloads2d agoHugging Face02MeghanaKap /yoruba_dataset_encodedgatedtext100K<n<1M0 likes963 downloads16d agoHugging Face03bytel0rd /yoruba_audio_translatedThis is a copy of odunola/Yoruba_translate_preprocessed, the only difference is, it's already splitted into train & test. Awesome credits to her, her license applies too. audiotranslation10K<n<100K3 likes490 downloads2y agoHugging Face04michsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes395 downloads1y agoHugging Face05voicedata /9jalingo-reviewed-yorubaaudio1K<n<10K0 likes219 downloads5d agoHugging Face06Mawube /s2tt-yoruba-englishaudio10K<n<100K0 likes216 downloads6mo agoHugging Face07FloatinggOnion /yoruba-cfm-latentstextn<1K0 likes182 downloads4mo agoHugging Face08MeghanaKap /yoruba_dataset_encoded_repo20to30text100K<n<1M0 likes144 downloads21d agoHugging Face09okezieowen /afrispeech_yorubaaudio10K<n<100K0 likes136 downloads1y agoHugging Face10voicedata /final_yorubagatedaudio100K<n<1M2 likes111 downloads17d agoHugging Face11adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes101 downloads13d agoHugging Face12EYEDOL /naija-voices-yoruba-split_2-3audio10K<n<100K0 likes100 downloads1y agoHugging Face130xnu /yoruba Yoruba Dataset The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Yoruba words. Data Types Synthetic sentences (400k): Basic vocabulary combinations Pattern variations (200k): Template-based grammatical structures Conversations (150k): Interactive dialogue examples Q&A pairs (100k): Knowledge and reasoning tasks Translations (80k): Yoruba-English bidirectional pairs Grammar examples (40k): Verb… See the full description on the dataset page: https://huggingface.co/datasets/0xnu/yoruba.tabular100K<n<1M0 likes97 downloads1y agoHugging Face14nolimitsxl /yfacc_yorubaaudio1K<n<10K0 likes80 downloads7d agoHugging Face15Arowoshola /yoruba-settext10K<n<100K2 likes77 downloads3y agoHugging Face16EYEDOL /naija-voices-yoruba-split_0-8audio10K<n<100K0 likes77 downloads1y agoHugging Face17EYEDOL /naija-voices-yoruba-split_0-6audio10K<n<100K0 likes76 downloads1y agoHugging Face18EYEDOL /naija-voices-yoruba-split_0-7audio10K<n<100K0 likes76 downloads1y agoHugging Face19EYEDOL /naija-voices-yoruba-split_2-7audio10K<n<100K0 likes74 downloads1y agoHugging Face20EYEDOL /naija-voices-yoruba-split_0-5audio10K<n<100K0 likes73 downloads1y agoHugging Face21EYEDOL /naija-voices-yoruba-split_2-0audio10K<n<100K0 likes73 downloads1y agoHugging Face22EYEDOL /naija-voices-yoruba-split_1-0audio10K<n<100K0 likes70 downloads1y agoHugging Face23EYEDOL /naija-voices-yoruba-split_2-2audio10K<n<100K0 likes70 downloads1y agoHugging Face24EYEDOL /naija-voices-yoruba-split_2-6audio10K<n<100K0 likes69 downloads1y agoHugging Face25EYEDOL /naija-voices-yoruba-split_0-4audio10K<n<100K0 likes66 downloads1y agoHugging Face26EYEDOL /naija-voices-yoruba-split_1-2audio10K<n<100K0 likes62 downloads1y agoHugging Face27EYEDOL /naija-voices-yoruba-split_1-6audio10K<n<100K0 likes61 downloads1y agoHugging Face28EYEDOL /naija-voices-yoruba-split_2-4audio10K<n<100K0 likes61 downloads1y agoHugging Face29DiCeyIII /Additional_Yoruba_Dataaudio1K<n<10K0 likes60 downloads2y agoHugging Face30michsethowusu /english-yoruba_sentence-pairs_mt560 English-Yoruba Parallel Dataset This dataset contains parallel sentences in English and Yoruba (Nigeria). Dataset Information Language Pair: English ↔ Yoruba Language Code: yor Country: Nigeria Original Source: OPUS MT560 Dataset Dataset Structure The dataset contains parallel sentences that can be used for: Machine translation training Cross-lingual NLP tasks Language model fine-tuning Citation If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-yoruba_sentence-pairs_mt560.text100K<n<1M0 likes57 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.