datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
naijaweb
Naijaweb Dataset 🇳🇬
Naijaweb is a dataset that contains over 270,000+ documents, totaling approximately 230 million GPT-2 tokens. The data was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string
language_probability
float64… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb.naija-customer-call-code-switch
Naija Customer-Call Code-Switch Corpus (Orinode-CCS)
Hand-written customer-service sentences with natural code-switching between Nigerian English and three indigenous Nigerian languages — Hausa, Yoruba, and Igbo. Covers 30+ business sectors typical of real customer-service calls in Nigeria.
Released by Orinode under CC-BY 4.0 to support research on multilingual ASR, NLU, and conversational AI for low-resource African languages.
Why this dataset exists
Global voice-AI… See the full description on the dataset page: https://huggingface.co/datasets/Orinode/naija-customer-call-code-switch.naijaweb-edu
Naijaweb Edu Dataset 🇳🇬
Naijaweb Edu is a subset of the naijaweb dataset with an educational score aboove 3 using the fineweb classifier. The initial fineweb dataset was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb-edu.naija-pidgin-health-qa-rivers-2026naijaweb-edu2
Naijaweb Edu2 Dataset 🇳🇬
Naijaweb Edu 2 is a subset of the naijaweb dataset with an educational score aboove 2 using the fineweb classifier. The initial fineweb dataset was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb-edu2.naija-agric-qa-synth
AgriPadi Synthetic SFT
Production-ready synthetic agriculture-advisory SFT data for West Africa
(English, Nigerian Pidgin, Hausa), generated by the AgriPadi pipelines.
Pipeline source (reproducible): https://github.com/ThatLinuxGuyYouKnow/agripadi_synth_pipeline
All rows are published in the regular chat format (messages as
user/assistant pairs). The raw-completion surface used for lm-eval
profiling is intentionally excluded.
Composition
6,701 training rows:… See the full description on the dataset page: https://huggingface.co/datasets/Alabi-Ayobami/naija-agric-qa-synth.
