datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.pi-dev-plugins
Pi Coding Agent Plugins
Dataset contains metadata for approximately 5500 Pi Coding Agent plugins gathered from https://pi.dev/packages on 2026-09-22. Row example:
{
"name": "pi-mcp-adapter",
"description": "MCP (Model Context Protocol) adapter extension for Pi coding agent",
"types": [
"extension"
],
"author": "nicopreme",
"downloads_monthly": 1013749,
"downloads_label": "1M/mo",
"published_label": "22h ago",
"published_ms": 1790009755031,
"detail_url":… See the full description on the dataset page: https://huggingface.co/datasets/kth8/pi-dev-plugins.pi_datasetMedical-Reasoning-Dataset-Nigerian-Pidgin
Medical Reasoning Dataset Nigerian Pidgin | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Medical-Reasoning-Dataset-Nigerian-Pidgin.Pidgin-QandA-data-samples
Pidgin Question-Answer Dataset (Sample)
Sample dataset: Nigerian Pidgin conversational Q&A for dialogue systems and language modeling
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin Question-Answer Dataset (Sample) is a conversational corpus containing 1,462 question-answer pairs entirely in Nigerian Pidgin English. Created by Bytte AI through AI chatbot interactions with human validation, this sample dataset supports dialogue… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-QandA-data-samples.digitize-pid-ner
Digitize-PID: Pipeline numbers (NER)
Note: I am not the author of this dataset
Named Entity Recognition dataset for extracting pipeline numbers from full text of P&ID
(Piping and Instrumentation Diagram) documents.
Dataset Details
Dataset Description
Pipeline numbers are structured identifiers in engineering documents:
Example Format: A-123-BC (3-5 segments with a separator such as -, , or _)
Use case: Automated extraction from P&ID document text
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-ner.pidgin_bank_dataset
Nigerian Pidgin Bank Customer Support Dataset
Overview
This dataset contains 150,000 single-turn customer support conversations for a Nigerian retail bank, written primarily in Nigerian Pidgin English (pcm) with code-switched English. Each example simulates a customer inquiry about common banking issues — such as failed transfers, POS/ATM dispense errors, USSD (*737#) banking, and internet banking login/password resets — paired with an assistant response grounded… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/pidgin_bank_dataset.Pidgin_Question-English_Answer_Dataset
Pidgin Question - English Answer Dataset (Sample)
Data Card v1.0
Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/
Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.mt5_nigerian_pidginDeveloped by rufatronics (Aga) Ahmad Garba Adamu
mt5_nigerian_pidgin
pidgin_clean
Pidgin Clean
Overview
pidgin_clean is a cleaned collection of short, two-turn conversational exchanges written in Nigerian Pidgin English, frequently code-mixed with standard English. It was assembled to support the development and fine-tuning of NLP models — particularly conversational/chat models — that need to understand and generate natural Nigerian Pidgin text.
Dataset Structure
The dataset contains 2,066 conversation records, stored as a… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/pidgin_clean.Pidgin-to-English-conversational-translations
Pidgin-to-English Translation Dataset (Sample)
Sample dataset: Nigerian Pidgin to English translation pairs for machine translation research
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin-to-English Translation Dataset (Sample) is a conversational-style parallel corpus containing 122 translation pairs from Nigerian Pidgin English to Standard English. Created by Bytte AI through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-to-English-conversational-translations.PID_patches_syn_v2orpheus-pidgin-cptyoruba_pidgin_agriculture_dataPID_patches_synpidgin-corpus-synthpidgin_csvpidgin_conversation_datasetpi-defense-experiment-dataset
Prompt Injection Defense Experiment Dataset
This dataset contains the processed data used in an experimental evaluation of prompt-injection defenses for instruction-following language models.
Repository: leinha/pi-defense-experiment-dataset
Intended use
Academic and experimental evaluation of prompt-injection defenses, including StruQ-like SFT, SecAlign-like DPO and Instruction-Hierarchy-like SFT scenarios.
Content warning
This dataset may contain… See the full description on the dataset page: https://huggingface.co/datasets/leinha/pi-defense-experiment-dataset.process-pid-extraction-dataset-100kpi-datasetPidginversion2pidgin0.2process-pid-extraction-small-100k
