datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NaijaS2ST
Multilingual Speech Dataset (Speech-to-Speech / Speech-to-Text Ready)
*The IWSLT shared task submission details and the test set are now available at IWSLT 2026 *
Dataset Summary
This dataset is a large-scale multilingual speech corpus curated for speech-to-speech translation, speech-to-text, and multilingual speech processing research.
The data is organized by language, speaker (user_id), and dataset split (train, dev), and includes rich acoustic and metadata… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/NaijaS2ST.naija3naija2naijavoices-dataset
Important Information: please be aware that this version of the dataset is huge (500+ GB) and can therefore be challenging to use. To alleviate this and facilitate adoption, we’ve provided a compressed version (84GB) here. We suggest using that instead if you have compute/storage constraints. They are both the exact same data.
Introduction
Welcome to the NaijaVoices dataset. The NaijaVoices dataset consists of 1,800 hours of authentic speech (from over 5,000 diverse speakers!)… See the full description on the dataset page: https://huggingface.co/datasets/naijavoices/naijavoices-dataset.African_voices_naija
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Nigerian Pidgin (pcm) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Nigerian Pidgin (pcm). This corpus is designed to accelerate the development of speech technology in African contexts, promoting… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_naija.naija-speech-afrispeech-ngmozilla_commonvoice_naijaHausa1_preprocessed_train_batch_1naija_englishafricanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
NaijaSenti-TwitterNaijaSenti is the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria — Hausa, Igbo, Nigerian-Pidgin, and Yorùbá — consisting of around 30,000 annotated tweets per language, including a significant fraction of code-mixed tweets.naijaweb
Naijaweb Dataset 🇳🇬
Naijaweb is a dataset that contains over 270,000+ documents, totaling approximately 230 million GPT-2 tokens. The data was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string
language_probability
float64… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb.naija-voices-hausa-split_0-1NaijaRC
Dataset Card for NaijaRC
Dataset Summary
Languages
There are 3 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data = load_dataset('Davlan/NaijaRC', 'yor')
# Please, specify the language code
# A data point example is below:
@article{aremu2023naijarc,
title={Naijarc: A multi-choice reading comprehension dataset for nigerian languages}… See the full description on the dataset page: https://huggingface.co/datasets/Davlan/NaijaRC.naija-voices-igbo-split_1-1naija-voices-yoruba-split_2-3naijanaija_sftkiri-naija-voice-datasetnaija-voices-hausa-split_0-6naija-voices-yoruba-split_0-6naija-voices-hausa-split_0-4naija-voices-hausa-split_2-4naija-voices-hausa-split_0-7naija-voices-yoruba-split_0-8naija-voices-hausa-split_0-5naija-voices-yoruba-split_0-5naija-voices-yoruba-split_0-7naija-voices-igbo-split_2-1naija-voices-yoruba-split_2-7naija-voices-yoruba-split_2-0
