datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/Sri1311/pid_lines_dataset.pid-icons-mergedpi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.final_pidginsample_pidppi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.9jalingo-reviewed-pidginnigerian-pidgin-1.0
Language:
- Nigerian Pidgin English (West African Pidgin variant)
Dataset Description
Dataset Summary
The Nigerian Pidgin ASR dataset (v1.0) is the first publicly released speech-to-text corpus for Nigerian Pidgin English, a widely spoken lingua franca across Nigeria and West Africa. This dataset comprises over 3,000 audio recordings paired with sentence-level transcriptions, recorded by native speakers across different genders and age groups. It is tailored for… See the full description on the dataset page: https://huggingface.co/datasets/asr-nigerian-pidgin/nigerian-pidgin-1.0.digitize-pid-yolo
Digitize-PID (Symbols only), YOLO format
Note: I am not the author of this dataset
An annotated synthetic dataset, Dataset-P&ID, of 500 P&IDs with incorporates different
types of noise and complex symbols. This dataset contains only the symbols, i.e., under
the object detection task.
Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021). Digitize-PID: Automatic Digitization
of Piping and Instrumentation Diagrams. ArXiv, abs/2109.03794.
Original data repository:… See the full description on the dataset page: https://huggingface.co/datasets/SherifAhmed/digitize-pid-yolo.Pidgin_ASR_Dataset_Combined
Naija-ASR-Corpus v1.0 (NAC-v1.0)
A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM)
📌 Dataset Summary
Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC).
The NAC Team processed the original long-form recordings by:
Segmenting the audio into sentence-level clips.
Transcribing/Aligning the text to create paired audio-text data suitable for ASR training.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/timniel/Pidgin_ASR_Dataset_Combined.Prompt_Injection_PIDSpi-dev-plugins
Pi Coding Agent Plugins
Dataset contains metadata for approximately 5500 Pi Coding Agent plugins gathered from https://pi.dev/packages on 2026-09-22. Row example:
{
"name": "pi-mcp-adapter",
"description": "MCP (Model Context Protocol) adapter extension for Pi coding agent",
"types": [
"extension"
],
"author": "nicopreme",
"downloads_monthly": 1013749,
"downloads_label": "1M/mo",
"published_label": "22h ago",
"published_ms": 1790009755031,
"detail_url":… See the full description on the dataset page: https://huggingface.co/datasets/kth8/pi-dev-plugins.Pidgin_ASR_Dataset_Combined
Naija-ASR-Corpus v1.0 (NAC-v1.0)
A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM)
📌 Dataset Summary
Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC).
The NAC Team processed the original long-form recordings by:
Segmenting the audio into sentence-level clips.
Transcribing/Aligning the text to create paired audio-text data suitable for ASR training.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/AnonXx/Pidgin_ASR_Dataset_Combined.pidgin-asr-combined
Pidgin ASR Combined
A unified Nigerian Pidgin English speech-to-text dataset that combines
publicly available Pidgin ASR sources into a single train / validation /
test setup with a consistent schema. Built for fine-tuning Whisper-family
models on Nigerian Pidgin (Naija, pcm).
~8.6 hours, 4,278 clips, 10 source speakers, 16 kHz mono WAV.
Used to train michaelodafe/whisper-pidgin-v1
(21.37% WER on the test split, beating the published Wav2Vec2-XLSR-53
baseline by 8.2 pp).… See the full description on the dataset page: https://huggingface.co/datasets/michaelodafe/pidgin-asr-combined.english-nigerian-pidgin_sentence-pairs_mt560
English-Nigerian Pidgin Parallel Dataset
This dataset contains parallel sentences in English and Nigerian Pidgin (Nigeria).
Dataset Information
Language Pair: English ↔ Nigerian Pidgin
Language Code: pcm
Country: Nigeria
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-nigerian-pidgin_sentence-pairs_mt560.pidginData
Naija-ASR-Corpus v2.0 (NAC-v2.0)
A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM)
📌 Dataset Summary
Naija-ASR-Corpus (NAC-v2.0) is a speech dataset derived from the Universal
Dependencies Naija Spoken Corpus (UD_Naija-NSC).
The NAC Team processed the original long-form recordings by:
Segmenting the audio into sentence-level clips.
Aligning each clip to its transcript (text_ortho) from the CoNLL-U source.
Tagging each sample… See the full description on the dataset page: https://huggingface.co/datasets/voicedata/pidginData.pi_datasetMedical-Reasoning-Dataset-Nigerian-Pidgin
Medical Reasoning Dataset Nigerian Pidgin | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Medical-Reasoning-Dataset-Nigerian-Pidgin.piqa_yoruba_pidgin
Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin
Dataset Summary
This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures.
It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.Pidgin-QandA-data-samples
Pidgin Question-Answer Dataset (Sample)
Sample dataset: Nigerian Pidgin conversational Q&A for dialogue systems and language modeling
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin Question-Answer Dataset (Sample) is a conversational corpus containing 1,462 question-answer pairs entirely in Nigerian Pidgin English. Created by Bytte AI through AI chatbot interactions with human validation, this sample dataset supports dialogue… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-QandA-data-samples.BBC_Igbo-Pidgin_Gold-Standard_NLP_Corpus
BBC Igbo–Pidgin Gold-Standard NLP Corpus (Sample)
Sample dataset: High-quality annotated data for Nigerian Igbo and Pidgin English NLP
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The BBC Igbo–Pidgin Gold-Standard NLP Corpus (Sample) is a meticulously curated collection of professionally annotated text data designed to advance natural language processing for Nigerian languages. Created by Bytte AI, this sample corpus addresses the… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/BBC_Igbo-Pidgin_Gold-Standard_NLP_Corpus.en-pidgin-all-audio-57minsnigerian-pidgin-speech
Nigerian Pidgin Audio + Text Dataset for Whisper Fine-tuning
Nigerian Pidgin speech dataset for Whisper fine-tuning
Dataset Summary
This dataset contains audio recordings and transcriptions in Nigerian Pidgin English, designed for fine-tuning speech recognition models, particularly OpenAI's Whisper.
Dataset Structure
Train Split: 65 samples
Test Split: 8 samples
Total Duration: 0.0 hours (estimated)
Average Duration: 2.4 seconds per sample
Sample Rate: 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/Rexe/nigerian-pidgin-speech.celeb_pidgindigitize-pid-ner
Digitize-PID: Pipeline numbers (NER)
Note: I am not the author of this dataset
Named Entity Recognition dataset for extracting pipeline numbers from full text of P&ID
(Piping and Instrumentation Diagram) documents.
Dataset Details
Dataset Description
Pipeline numbers are structured identifiers in engineering documents:
Example Format: A-123-BC (3-5 segments with a separator such as -, , or _)
Use case: Automated extraction from P&ID document text
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-ner.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line segments as [x1… See the full description on the dataset page: https://huggingface.co/datasets/SherifAhmed/pid_lines_dataset.english-cameroon-pidgin_sentence-pairs_mt560
English-Cameroon Pidgin Parallel Dataset
This dataset contains parallel sentences in English and Cameroon Pidgin (Cameroon).
Dataset Information
Language Pair: English ↔ Cameroon Pidgin
Language Code: wes
Country: Cameroon
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-cameroon-pidgin_sentence-pairs_mt560.eng-pidgin-tts-july
eng-pidgin-tts-july
Local TTS dataset uploaded from this repository.
Source
Data dir: /home/ml/workspaces/kashish/data-creation/data-prep-eleven-labs/output/combined_july7
CSV: /home/ml/workspaces/kashish/data-creation/data-prep-eleven-labs/output/combined_july7/data.csv
Transcript column: transcript
Speaker column: speaker
Audio sample rate: 24000
Columns
Column
Description
audio
Local WAV file uploaded as Hugging Face audio feature… See the full description on the dataset page: https://huggingface.co/datasets/vaghawan/eng-pidgin-tts-july.pidgin_bank_dataset
Nigerian Pidgin Bank Customer Support Dataset
Overview
This dataset contains 150,000 single-turn customer support conversations for a Nigerian retail bank, written primarily in Nigerian Pidgin English (pcm) with code-switched English. Each example simulates a customer inquiry about common banking issues — such as failed transfers, POS/ATM dispense errors, USSD (*737#) banking, and internet banking login/password resets — paired with an assistant response grounded… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/pidgin_bank_dataset.naija-pidgin-health-qa-rivers-2026
