CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sri1311 /pid_lines_dataset P&ID Line Detection Dataset This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams) with line segment annotations for line detection and segmentation tasks. Dataset Structure Each sample contains: file_name: Image filename source_image_idx: Index of the original P&ID image crop_idx: Index of this crop from the source image width: Crop width in pixels height: Crop height in pixels lines: Dictionary with: segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/Sri1311/pid_lines_dataset.imageimage-segmentation10K<n<100K1 likes601 downloads2mo agoHugging Face02maxs-m87 /pid-icons-mergedimage10K<n<100K0 likes303 downloads7mo agoHugging Face03badlogicgames /pi-diff-review Coding agent session traces for badlogicgames/pi-diff-review This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.tabulartext-generationn<1K8 likes238 downloads6mo agoHugging Face04voicedata /final_pidgingatedaudio100K<n<1M0 likes206 downloads16d agoHugging Face05LegPiece /sample_pidptext10K<n<100K0 likes150 downloads1y agoHugging Face06cfahlgren1 /pi-diff-review Coding agent session traces for badlogicgames/pi-diff-review This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.tabulartext-generationn<1K0 likes146 downloads6mo agoHugging Face07voicedata /9jalingo-reviewed-pidginaudion<1K0 likes133 downloads7d agoHugging Face08asr-nigerian-pidgin /nigerian-pidgin-1.0 Language: - Nigerian Pidgin English (West African Pidgin variant) Dataset Description Dataset Summary The Nigerian Pidgin ASR dataset (v1.0) is the first publicly released speech-to-text corpus for Nigerian Pidgin English, a widely spoken lingua franca across Nigeria and West Africa. This dataset comprises over 3,000 audio recordings paired with sentence-level transcriptions, recorded by native speakers across different genders and age groups. It is tailored for… See the full description on the dataset page: https://huggingface.co/datasets/asr-nigerian-pidgin/nigerian-pidgin-1.0.audioautomatic-speech-recognition1K<n<10K2 likes118 downloads1y agoHugging Face09SherifAhmed /digitize-pid-yolo Digitize-PID (Symbols only), YOLO format Note: I am not the author of this dataset An annotated synthetic dataset, Dataset-P&ID, of 500 P&IDs with incorporates different types of noise and complex symbols. This dataset contains only the symbols, i.e., under the object detection task. Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021). Digitize-PID: Automatic Digitization of Piping and Instrumentation Diagrams. ArXiv, abs/2109.03794. Original data repository:… See the full description on the dataset page: https://huggingface.co/datasets/SherifAhmed/digitize-pid-yolo.imageobject-detection10K<n<100K0 likes118 downloads3mo agoHugging Face10timniel /Pidgin_ASR_Dataset_Combined Naija-ASR-Corpus v1.0 (NAC-v1.0) A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM) 📌 Dataset Summary Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC). The NAC Team processed the original long-form recordings by: Segmenting the audio into sentence-level clips. Transcribing/Aligning the text to create paired audio-text data suitable for ASR training. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/timniel/Pidgin_ASR_Dataset_Combined.audioautomatic-speech-recognition1K<n<10K0 likes117 downloads10mo agoHugging Face11MuhammadAnas1657 /Prompt_Injection_PIDStext100K<n<1M1 likes116 downloads22d agoHugging Face12kth8 /pi-dev-plugins Pi Coding Agent Plugins Dataset contains metadata for approximately 5500 Pi Coding Agent plugins gathered from https://pi.dev/packages on 2026-09-22. Row example: { "name": "pi-mcp-adapter", "description": "MCP (Model Context Protocol) adapter extension for Pi coding agent", "types": [ "extension" ], "author": "nicopreme", "downloads_monthly": 1013749, "downloads_label": "1M/mo", "published_label": "22h ago", "published_ms": 1790009755031, "detail_url":… See the full description on the dataset page: https://huggingface.co/datasets/kth8/pi-dev-plugins.tabular1K<n<10K0 likes109 downloads2h agoHugging Face13AnonXx /Pidgin_ASR_Dataset_Combined Naija-ASR-Corpus v1.0 (NAC-v1.0) A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM) 📌 Dataset Summary Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC). The NAC Team processed the original long-form recordings by: Segmenting the audio into sentence-level clips. Transcribing/Aligning the text to create paired audio-text data suitable for ASR training. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/AnonXx/Pidgin_ASR_Dataset_Combined.audioautomatic-speech-recognition1K<n<10K1 likes107 downloads6mo agoHugging Face14michaelodafe /pidgin-asr-combined Pidgin ASR Combined A unified Nigerian Pidgin English speech-to-text dataset that combines publicly available Pidgin ASR sources into a single train / validation / test setup with a consistent schema. Built for fine-tuning Whisper-family models on Nigerian Pidgin (Naija, pcm). ~8.6 hours, 4,278 clips, 10 source speakers, 16 kHz mono WAV. Used to train michaelodafe/whisper-pidgin-v1 (21.37% WER on the test split, beating the published Wav2Vec2-XLSR-53 baseline by 8.2 pp).… See the full description on the dataset page: https://huggingface.co/datasets/michaelodafe/pidgin-asr-combined.audioautomatic-speech-recognition1K<n<10K1 likes91 downloads4mo agoHugging Face15michsethowusu /english-nigerian-pidgin_sentence-pairs_mt560 English-Nigerian Pidgin Parallel Dataset This dataset contains parallel sentences in English and Nigerian Pidgin (Nigeria). Dataset Information Language Pair: English ↔ Nigerian Pidgin Language Code: pcm Country: Nigeria Original Source: OPUS MT560 Dataset Dataset Structure The dataset contains parallel sentences that can be used for: Machine translation training Cross-lingual NLP tasks Language model fine-tuning Citation If you use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-nigerian-pidgin_sentence-pairs_mt560.text10K<n<100K1 likes84 downloads1y agoHugging Face16voicedata /pidginData Naija-ASR-Corpus v2.0 (NAC-v2.0) A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM) 📌 Dataset Summary Naija-ASR-Corpus (NAC-v2.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC). The NAC Team processed the original long-form recordings by: Segmenting the audio into sentence-level clips. Aligning each clip to its transcript (text_ortho) from the CoNLL-U source. Tagging each sample… See the full description on the dataset page: https://huggingface.co/datasets/voicedata/pidginData.audioautomatic-speech-recognition1K<n<10K0 likes72 downloads23d agoHugging Face17raymondt /pi_datasettext10K<n<100K0 likes63 downloads1y agoHugging Face18electricsheepafrica /Medical-Reasoning-Dataset-Nigerian-Pidgin Medical Reasoning Dataset Nigerian Pidgin | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Medical-Reasoning-Dataset-Nigerian-Pidgin.texttabular-classification10K<n<100K0 likes46 downloads1mo agoHugging Face19taresco /piqa_yoruba_pidgin Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin Dataset Summary This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures. It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios. The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.textquestion-answeringn<1K2 likes45 downloads10mo agoHugging Face20Bytte-AI /Pidgin-QandA-data-samples Pidgin Question-Answer Dataset (Sample) Sample dataset: Nigerian Pidgin conversational Q&A for dialogue systems and language modeling 🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact 📋 Overview The Pidgin Question-Answer Dataset (Sample) is a conversational corpus containing 1,462 question-answer pairs entirely in Nigerian Pidgin English. Created by Bytte AI through AI chatbot interactions with human validation, this sample dataset supports dialogue… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-QandA-data-samples.texttext-classification1K<n<10K0 likes44 downloads8mo agoHugging Face21Bytte-AI /BBC_Igbo-Pidgin_Gold-Standard_NLP_Corpus BBC Igbo–Pidgin Gold-Standard NLP Corpus (Sample) Sample dataset: High-quality annotated data for Nigerian Igbo and Pidgin English NLP 🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact 📋 Overview The BBC Igbo–Pidgin Gold-Standard NLP Corpus (Sample) is a meticulously curated collection of professionally annotated text data designed to advance natural language processing for Nigerian languages. Created by Bytte AI, this sample corpus addresses the… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/BBC_Igbo-Pidgin_Gold-Standard_NLP_Corpus.texttext-generationn<1K1 likes42 downloads8mo agoHugging Face22vaghawan /en-pidgin-all-audio-57minsaudion<1K0 likes37 downloads1mo agoHugging Face23Rexe /nigerian-pidgin-speech Nigerian Pidgin Audio + Text Dataset for Whisper Fine-tuning Nigerian Pidgin speech dataset for Whisper fine-tuning Dataset Summary This dataset contains audio recordings and transcriptions in Nigerian Pidgin English, designed for fine-tuning speech recognition models, particularly OpenAI's Whisper. Dataset Structure Train Split: 65 samples Test Split: 8 samples Total Duration: 0.0 hours (estimated) Average Duration: 2.4 seconds per sample Sample Rate: 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/Rexe/nigerian-pidgin-speech.audioautomatic-speech-recognitionn<1K1 likes36 downloads1y agoHugging Face24okezieowen /celeb_pidginaudio1K<n<10K0 likes34 downloads1y agoHugging Face25hamzas /digitize-pid-ner Digitize-PID: Pipeline numbers (NER) Note: I am not the author of this dataset Named Entity Recognition dataset for extracting pipeline numbers from full text of P&ID (Piping and Instrumentation Diagram) documents. Dataset Details Dataset Description Pipeline numbers are structured identifiers in engineering documents: Example Format: A-123-BC (3-5 segments with a separator such as -, , or _) Use case: Automated extraction from P&ID document text Domain:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-ner.texttoken-classificationn<1K0 likes32 downloads11mo agoHugging Face26SherifAhmed /pid_lines_dataset P&ID Line Detection Dataset This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams) with line segment annotations for line detection and segmentation tasks. Dataset Structure Each sample contains: file_name: Image filename source_image_idx: Index of the original P&ID image crop_idx: Index of this crop from the source image width: Crop width in pixels height: Crop height in pixels lines: Dictionary with: segments: List of line segments as [x1… See the full description on the dataset page: https://huggingface.co/datasets/SherifAhmed/pid_lines_dataset.imageimage-segmentation10K<n<100K0 likes32 downloads5mo agoHugging Face27michsethowusu /english-cameroon-pidgin_sentence-pairs_mt560 English-Cameroon Pidgin Parallel Dataset This dataset contains parallel sentences in English and Cameroon Pidgin (Cameroon). Dataset Information Language Pair: English ↔ Cameroon Pidgin Language Code: wes Country: Cameroon Original Source: OPUS MT560 Dataset Dataset Structure The dataset contains parallel sentences that can be used for: Machine translation training Cross-lingual NLP tasks Language model fine-tuning Citation If you use this… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-cameroon-pidgin_sentence-pairs_mt560.text10K<n<100K1 likes28 downloads1y agoHugging Face28vaghawan /eng-pidgin-tts-july eng-pidgin-tts-july Local TTS dataset uploaded from this repository. Source Data dir: /home/ml/workspaces/kashish/data-creation/data-prep-eleven-labs/output/combined_july7 CSV: /home/ml/workspaces/kashish/data-creation/data-prep-eleven-labs/output/combined_july7/data.csv Transcript column: transcript Speaker column: speaker Audio sample rate: 24000 Columns Column Description audio Local WAV file uploaded as Hugging Face audio feature… See the full description on the dataset page: https://huggingface.co/datasets/vaghawan/eng-pidgin-tts-july.audio1K<n<10K0 likes28 downloads3mo agoHugging Face29Ephraimmm /pidgin_bank_dataset Nigerian Pidgin Bank Customer Support Dataset Overview This dataset contains 150,000 single-turn customer support conversations for a Nigerian retail bank, written primarily in Nigerian Pidgin English (pcm) with code-switched English. Each example simulates a customer inquiry about common banking issues — such as failed transfers, POS/ATM dispense errors, USSD (*737#) banking, and internet banking login/password resets — paired with an assistant response grounded… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/pidgin_bank_dataset.texttext-generation100K<n<1M0 likes27 downloads3mo agoHugging Face30Charley890 /naija-pidgin-health-qa-rivers-2026textquestion-answeringn<1K0 likes27 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.