datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
African_voices_igbo
🇳🇬 WaZoBiaSpeech: 500+ Hour Igbo (ibo) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Igbo (ibo). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_igbo.igbo_tts_normalizeds2tt-igbo-englishclean_igbo_datasetigbo-dict-expansion-16khzBased on: https://huggingface.co/datasets/nkowaokwu/ibo-dict-expansion
The original audios were converted to 16kHz WAV.
Citation
If you want to cite this dataset you can use this:
@misc{igbo-dict-expansion-16khz,
title={Igbo dataset},
author={Jimenez, David},
howpublished={\url{https://huggingface.co/datasets/deepdml/igbo-dict-expansion-16khz}},
year={2025}
}
naija-voices-igbo-split_1-1igbo-dict-expansionigbo-dict-16khzigbo-dict is an Igbo text-audio dataset that includes the following:
25,500 single word audio recordings for each dialectal word variation
25,000 single Igbo sentence audio recording for each Igbo-English sentence pairing
The original audios were converted to 16kHz WAV.
Referenced in The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment
Citation
If you want to cite this dataset you can use this:
@misc{igbo-dict-16khz,
title={Igbo… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/igbo-dict-16khz.igbo-speech
Igbo Speech Dataset
A protocol sample. 5 contributors, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment. Small by design — see What this is for below before downloading.
Hours
0.37
Clips
36
Speakers
5
Origin varieties
2
Languages
1
Configs
1
Speaker metadata
origin region / variety, mother tongue, gender, device, OS, recording environment
Audio… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/igbo-speech.naija-voices-igbo-split_2-1naija-voices-igbo-split_2-9igbo-dictomniASR-igbo-blindspots
omniASR Igbo Blind Spot Dataset
Research Questions
This dataset investigates three interrelated questions about multilingual ASR performance on tonal languages:
Operational Definition: What does "language support" mean when a model lists 1,600+ languages? Does coverage imply functional accuracy on linguistically meaningful distinctions?
Diagnostic Validity: Can tonal diacritic preservation serve as a diagnostic for acoustic competence vs. orthographic pattern matching… See the full description on the dataset page: https://huggingface.co/datasets/Chiz/omniASR-igbo-blindspots.naija-voices-igbo-split_0-9igbo_sync_rawnaija-voices-igbo-split_0-7naija-voices-igbo-split_0-4naija-voices-igbo-split_0-6naija-voices-igbo-split_0-82nd_igbosyncorp-igbo-asr-benchmarkigbo_sync_processedigbo_filtered_dataset_2final_igbonaija-voices-igbo-split_0-5naija-voices-igbo-split_0-2naija-voices-igbo-split_0-0naija-voices-igbo-split_0-4naija-voices-igbo-split_1-99javoice-igbo
9jaVoice Consent-1
169 clips. 30.5 minutes, which is 0.5 hours. 11 speakers. Igbo. FLAC, 48 kHz, mono, 16-bit. CC BY-NC 4.0.
Read-aloud speech, recorded by paid contributors on their own phones. Use it for evaluation, for fine-tuning, and as a reference set when you want to find out whether a model handles Nigerian speech at all. It is too small to pretrain on and we are not going to pretend otherwise.
Every clip here carries its own consent record. The contributor ticked an… See the full description on the dataset page: https://huggingface.co/datasets/9jatesters/9javoice-igbo.naija-voices-igbo-split_0-6
