datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
African_voices_naija
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Nigerian Pidgin (pcm) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Nigerian Pidgin (pcm). This corpus is designed to accelerate the development of speech technology in African contexts, promoting… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_naija.africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
voiceafrica-datasets
VoiceAfrica Dataset
Introduction
Welcome to the VoiceAfrica dataset. VoiceAfrica is a Lanfrica–Meta collaboration (code-named VoiceAfrica 1) that set out to create authentic, conversational speech and expert-curated transcriptions for 11 under-represented African languages. The dataset contains ~125 hours of speech (about 10 hours per language) across 16,383 audio samples from 118 speakers in four countries.
Unlike read-speech corpora, VoiceAfrica uses a natural… See the full description on the dataset page: https://huggingface.co/datasets/naijavoices/voiceafrica-datasets.
