naijavoices/voiceafrica-datasets
VoiceAfrica Dataset Introduction Welcome to the VoiceAfrica dataset. VoiceAfrica is a Lanfrica–Meta collaboration (code-named VoiceAfrica 1) that set out to create authentic, conversational speech and expert-curated transcriptions for 11 under-represented African languages. The dataset contains ~125 hours of speech (about 10 hours per language) across 16,383 audio samples from 118 speakers in four countries. Unlike read-speech corpora, VoiceAfrica uses a natural… See the full description on the dataset page: https://huggingface.co/datasets/naijavoices/voiceafrica-datasets.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face