datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
NaijaRC
Dataset Card for NaijaRC
Dataset Summary
Languages
There are 3 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data = load_dataset('Davlan/NaijaRC', 'yor')
# Please, specify the language code
# A data point example is below:
@article{aremu2023naijarc,
title={Naijarc: A multi-choice reading comprehension dataset for nigerian languages}… See the full description on the dataset page: https://huggingface.co/datasets/Davlan/NaijaRC.NaijaRC
Dataset Card for afrixnli
Dataset Summary
Languages
There are 3 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data = load_dataset('aremuadeolajr/NaijaRC', 'yor')
# Please, specify the language code
# A data point example is below:
naijavoices-dataset-compressed
Introduction
Welcome to the NaijaVoices dataset. The NaijaVoices dataset consists of 1,800 hours of authentic speech (from over 5,000 diverse speakers!) and expert curated text in Igbo, Hausa, and Yoruba. ~600 hours for each of the three languages. It also boasts of adequate female representation and balanced age-range distribution (young to old speakers). For more about the dataset info visit our website: https://naijavoices.com/. By using this dataset, you acknowledge reading… See the full description on the dataset page: https://huggingface.co/datasets/naijavoices/naijavoices-dataset-compressed.naija-pidgin-health-qa-rivers-2026NaijaHate
Dataset Card for NaijaHate
NaijaHate is a hate speech dataset tailored to the Nigerian context. It contains 35,973 annotated Nigerian tweets, including 29,999 tweets randomly sampled from Nigerian Twitter. For a complete description of the data, please refer to the reference paper.
Source Data
This dataset was sourced from a large Twitter dataset of 2.2 billion tweets posted between March 2007 and July 2023 and forming the timelines of 2.8 million Twitter users with a… See the full description on the dataset page: https://huggingface.co/datasets/worldbank/NaijaHate.
