WAZOBIALABS/nigerian-pidgin-voice-text
Wazobia Labs — Nigerian Pidgin Emotion & Sentiment Dataset Version: v0.8 — May 2026 Entries: 550 annotated entries Builder: Wazobia Labs License: CC-BY-4.0 — commercial use permitted with attribution Contact: wazobialabs@gmail.com Language: Nigerian Pidgin (Naija) — ISO 639-3: pcm What This Is The first commercially licensed Nigerian Pidgin emotion and sentiment dataset — built from lived native speaker knowledge, not translated from English, not scraped from… See the full description on the dataset page: https://huggingface.co/datasets/WAZOBIALABS/nigerian-pidgin-voice-text.
Wazobia Labs — Nigerian Pidgin Emotion & Sentiment Dataset
Version: v0.8 — May 2026 Entries: 550 annotated entries Builder: Wazobia Labs License: CC-BY-4.0 — commercial use permitted with attribution Contact: wazobialabs@gmail.com Language: Nigerian Pidgin (Naija) — ISO 639-3: pcm
What This Is
The first commercially licensed Nigerian Pidgin emotion and sentiment dataset — built from lived native speaker knowledge, not translated from English, not scraped from news broadcasts, not built for academic publication.
Built for the product team, not the paper.
Nigerian Pidgin is spoken by over 100 million people across Nigeria and the Nigerian diaspora. Every AI system currently deployed for Nigerian users — health chatbots, fintech customer service, voice assistants — is operating without any verified understanding of how Nigerian Pidgin speakers actually communicate. There is no evaluation benchmark. There is no culturally grounded emotion taxonomy. There is no sarcasm corpus.
Wazobia Labs builds what they deliberately left out.
Dataset Summary
The 16-Category Emotion Taxonomy
This dataset introduces the Wazobia Labs Nigerian Pidgin emotion taxonomy — 16 categories, four with no equivalent in any existing NLP framework:
Category Distribution
Why This Dataset Exists
The BBC Pidgin Problem
The most cited Nigerian Pidgin NLP resource is the BBC Pidgin corpus — compiled from formal news broadcasts. It has four sentiment labels. It contains no sarcasm pairs, almost no health language, and skews male.
A model trained on BBC Pidgin data cannot:
- Detect that "You don try well well" is sarcasm when delivered deadpan
- Classify "I no fit shout" as hustle_fatigue rather than generic negativity
- Read "E don do, I just dey manage" as clinical resignation rather than neutral filler
- Understand that "I dey my lane" is forming — performed indifference — not genuine contentment
Wazobia Labs builds the dataset that fills these gaps.
The Sarcasm Gap
Nigerian sarcasm is delivered deadpan — prosodically identical to sincere speech. The exact same phrase can mean the opposite depending entirely on cultural context.
This dataset contains 28 complete sarcasm pairs: the same Pidgin phrase annotated twice — once sincere, once sarcastic — with annotator notes documenting the contextual conditions that determine which reading applies.
The Health Domain
100 health domain entries capture how Nigerian patients communicate about their health in natural Pidgin. This matters because:
"E don do, I just dey manage" — AI reads neutral. It means a patient has given up trying to get better. That is a clinical signal.
"My body no dey again oh" — AI reads negative/frustration. It means physical depletion so severe the speaker cannot function normally.
Health AI deployed for Nigerian users without this layer of understanding will miss the moments that matter most.
Data Fields
Each entry contains 15 fields:
Annotation Methodology
Cultural Authority
Every annotation decision in this dataset was made by a native speaker with lived fluency across three Nigerian Pidgin registers:
- Warri Pidgin — the original creole form, Delta State origin
- Eastern Nigerian Pidgin — Owerri and Aba variants
- Lagos Pidgin — the urban cosmopolitan form
The lead annotator is an Igbo woman born in Aba, raised in Warri, educated in Owerri, living in Lagos. This trajectory is the methodology. Cultural authority — not institutional backing — is the foundation.
Bottom-Up Taxonomy Construction
Standard NLP emotion taxonomies are constructed top-down from existing psychological literature. This approach fails for Nigerian Pidgin because the psychological literature was not built on Nigerian speakers.
The Wazobia Labs taxonomy was constructed bottom-up: emotion categories were identified from observed native speaker expression in natural Pidgin communication, then formalised as annotation labels. Only categories that (a) cannot be accurately represented by existing taxonomy labels, (b) appear with sufficient frequency in natural Pidgin, and (c) can be defined clearly enough for consistent inter-annotator application were included.
Sarcasm Pair Protocol
Each sarcastic entry is paired with its sincere twin — the same phrase annotated as if spoken sincerely. Both entries carry prosody_match: matched because Nigerian sarcasm is prosodically identical to sincere speech. The sarcastic entry is additionally flagged sarcasm_flag: yes with annotator notes documenting the contextual conditions for the sarcastic reading.
Companion Datasets
[WAZOBIALABS/nigerian-pidgin-eval](https://huggingface.co/datasets/WAZOBIALABS/nigerian-pidgin-eval) Gold-standard evaluation set — v0.3 — 253 entries — all 16 categories at minimum 15 entries — 28 sarcasm pairs — 40 health entries
[WAZOBIALABS/igbo-voice-text](https://huggingface.co/datasets/WAZOBIALABS/igbo-voice-text) Igbo emotion dataset — v0.1 — 50 entries — 14 Igbo-specific emotion categories
Version History
Roadmap
Citation
@dataset{wazobia_labs_pidgin_2026,
author = {Okoye, Stephanie Nkemjika},
title = {Wazobia Labs Nigerian Pidgin Emotion and Sentiment Dataset},
year = {2026},
version = {0.8.0},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/WAZOBIALABS/nigerian-pidgin-voice-text},
license = {CC-BY-4.0},
note = {First commercially licensed Nigerian Pidgin emotion dataset with 16-category cultural taxonomy}
}Licensing
Published under CC-BY-4.0 — free to use for research, academic, and commercial purposes with attribution.
Enterprise licensing with support, update guarantees, version locking, and integration documentation available separately.
Contact: wazobialabs@gmail.com
About Wazobia Labs
Wazobia Labs builds African language AI infrastructure that does not exist but should. We identify the specific, high-value gaps in African language data that existing datasets leave open — then build exactly those gaps with commercial licensing, production-grade quality, and the cultural specificity that real AI products need.
Not another Twitter scrape. Not another scripted studio recording. We build what they deliberately left out.
