CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01michsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes395 downloads1y agoHugging Face02Professor /yoruba-speech-data Yoruba Speech Data (Pooled) A ~1,039-hour pooled Yoruba speech corpus, combining four independent sources into one consistently-formatted dataset for speech modeling (TTS / ASR). Built to give a Yoruba TTS finetune enough scale and diversity to move past what a single ~100-hour source can teach a model — more speakers, more domains, more of the language. Sources Source Clips Hours Style DSN African Voices (Data Science Nigeria) 144,628 309.0 h… See the full description on the dataset page: https://huggingface.co/datasets/Professor/yoruba-speech-data.text-to-speech100K<n<1M3 likes184 downloads1mo agoHugging Face03SilencioNetwork /yoruba-speech-transcribed Yoruba Spontaneous Speech, Transcribed — Silencio Spontaneous Yoruba with human-validated, fully tone-marked transcription and word-level alignment. 49 clips from 30 distinct speakers, 29 of them from Nigeria, across Ibadan, Lagos, Oyo and Nigerian Standard varieties. Transcripts keep the tonal diacritics and under-dots, and keep the Yoruba–English code-switching as it was spoken. Hours 0.53 Clips 49 Speakers 30 Countries 2 Speaker origin regions 7 Native… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/yoruba-speech-transcribed.audioautomatic-speech-recognitionn<1K0 likes52 downloads2d agoHugging Face04Kppwdfgu1 /yoruba-second-sbpn-demucs-20260826gated yoruba-second-sbpn-demucs-20260826 This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-second-sbpn-demucs-20260826.audioautomatic-speech-recognition1K<n<10K0 likes33 downloads27d agoHugging Face05Kppwdfgu1 /yoruba-datasold-sbpn-demucs-20260825gated yoruba-datasold-sbpn-demucs-20260825 This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-datasold-sbpn-demucs-20260825.audioautomatic-speech-recognitionn<1K0 likes28 downloads1mo agoHugging Face06Apakose-Ezekiel /farahan-yoruba-elder-corpus Apakose Ezekiel Imoleayo — Farahàn Appear. Be found. Lagos, Nigeria · UNILAG Yoruba Studies · Graduating 2030 What I Build Yoruba Oral Knowledge Corpus I document primary-source Yoruba knowledge directly from elder speakers in Lagos and southwest Nigeria. Structured interviews covering proverbs (Owe), oral history (Itan), praise poetry (Oriki), and cultural knowledge systems that no web scrape produces. What makes this corpus different:… See the full description on the dataset page: https://huggingface.co/datasets/Apakose-Ezekiel/farahan-yoruba-elder-corpus.audioautomatic-speech-recognitionn<1K0 likes25 downloads4mo agoHugging Face07Kppwdfgu1 /yoruba-bolanle-sbpn-demucs-20260825gated yoruba-bolanle-sbpn-demucs-20260825 This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-bolanle-sbpn-demucs-20260825.audioautomatic-speech-recognitionn<1K0 likes22 downloads29d agoHugging Face089jatesters /9javoice-yorubagated 9jaVoice Consent-1 153 clips. 29.4 minutes, which is 0.5 hours. 13 speakers. Yoruba. FLAC, 48 kHz, mono, 16-bit. CC BY-NC 4.0. Read-aloud speech, recorded by paid contributors on their own phones. Use it for evaluation, for fine-tuning, and as a reference set when you want to find out whether a model handles Nigerian speech at all. It is too small to pretrain on and we are not going to pretend otherwise. Every clip here carries its own consent record. The contributor ticked an… See the full description on the dataset page: https://huggingface.co/datasets/9jatesters/9javoice-yoruba.audioautomatic-speech-recognitionn<1K0 likes18 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.