datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PIDray
Dataset Card for pidray
PIDray is a large-scale dataset which covers various cases in real-world scenarios for prohibited item detection, especially for deliberately hidden items. The dataset contains 12 categories of prohibited items in 47, 677 X-ray images with high-quality annotated segmentation masks and bounding boxes.
This is a FiftyOne dataset with 9482 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/PIDray.digitize-pid-yolo
Digitize-PID (Symbols only), YOLO format
Note: I am not the author of this dataset
An annotated synthetic dataset, Dataset-P&ID, of 500 P&IDs with incorporates different
types of noise and complex symbols. This dataset contains only the symbols, i.e., under
the object detection task.
Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021). Digitize-PID: Automatic Digitization
of Piping and Instrumentation Diagrams. ArXiv, abs/2109.03794.
Original data repository:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-yolo.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/Sri1311/pid_lines_dataset.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line segments as [x1… See the full description on the dataset page: https://huggingface.co/datasets/prasatee/pid_lines_dataset.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/samiksha9874/pid_lines_dataset.PIDdatasetmetal_part_sort_v10_plus_extra_20260702
metal_part_sort_v10_plus_extra_20260702
Merged LeRobot v2.1 dataset for Unitree G1 + Inspire DFX metal part sorting.
Base dataset: PID0930/metal_part_sort_v10 @ 17fe43673ca30e840b1baa6e05c1f34b7dd43b91
Extra dataset: PID0930/metal_part_sort_v10_extra_20260701 @ 09c5a9113830443f99a8d91c11fd5754b238fc6a
Episodes: 329
Frames: 110966
Cameras: external, left wrist, right wrist, plus left_high/head slot retained for compatibility
GR00T training uses external as the logical head… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/metal_part_sort_v10_plus_extra_20260702.pid-icons-mergedpi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.final_pidgincloth_folding_v2_pidpsample_pidppi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.metal_part_sort_v6
metal_part_sort_v6
LeRobot v2.1 dataset for the Unitree G1 + Inspire DFX metal part sorting task.
Episodes: 39
Frames: 32,049
FPS: 30
Raw snapshot: /home/y-kosugi/datasets/metal_part_sort_v6_all_raw_20260624/metal_part_sort_v6
Camera streams are stored as:
observation.images.cam_left_high: robot head camera, recorded for audit only
observation.images.cam_external: external camera
observation.images.cam_left_wrist: left wrist camera
observation.images.cam_right_wrist: right… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/metal_part_sort_v6.9jalingo-reviewed-pidginmetal_part_sort_v3
metal_part_sort_v3
LeRobot v2.1 dataset for Unitree G1 + Inspire DFX metal part sorting.
Episodes: 95
Frames: 46748
Cameras: head, left wrist, right wrist
State/action: 26D, with hand dimensions 14:26
Hand-state convention: observation.state[t, 14:26] = action[t-1, 14:26] for t > 0; frame 0 uses same-frame hand action.
Failed recordings moved to trash were excluded before conversion.
metal_part_sort_v10
metal_part_sort_v10
LeRobot v2.1 dataset recorded on Unitree G1 + Inspire DFX for metal part sorting.
Episodes: 309
Frames: 101594
Cameras: external, left wrist, right wrist, head slot mapped to external for GR00T compatibility
State/action: 26 dims (left arm 7, right arm 7, left hand 6, right hand 6)
nigerian-pidgin-1.0
Language:
- Nigerian Pidgin English (West African Pidgin variant)
Dataset Description
Dataset Summary
The Nigerian Pidgin ASR dataset (v1.0) is the first publicly released speech-to-text corpus for Nigerian Pidgin English, a widely spoken lingua franca across Nigeria and West Africa. This dataset comprises over 3,000 audio recordings paired with sentence-level transcriptions, recorded by native speakers across different genders and age groups. It is tailored for… See the full description on the dataset page: https://huggingface.co/datasets/asr-nigerian-pidgin/nigerian-pidgin-1.0.digitize-pid-yolo
Digitize-PID (Symbols only), YOLO format
Note: I am not the author of this dataset
An annotated synthetic dataset, Dataset-P&ID, of 500 P&IDs with incorporates different
types of noise and complex symbols. This dataset contains only the symbols, i.e., under
the object detection task.
Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021). Digitize-PID: Automatic Digitization
of Piping and Instrumentation Diagrams. ArXiv, abs/2109.03794.
Original data repository:… See the full description on the dataset page: https://huggingface.co/datasets/SherifAhmed/digitize-pid-yolo.Pidgin_ASR_Dataset_Combined
Naija-ASR-Corpus v1.0 (NAC-v1.0)
A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM)
📌 Dataset Summary
Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC).
The NAC Team processed the original long-form recordings by:
Segmenting the audio into sentence-level clips.
Transcribing/Aligning the text to create paired audio-text data suitable for ASR training.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/timniel/Pidgin_ASR_Dataset_Combined.Prompt_Injection_PIDSPickPlaceCube_right_arm_pidpmetal_part_sort_v10_extra_20260701
metal_part_sort_v10_extra_20260701
Extra LeRobot v2.1 dataset recorded on Unitree G1 + Inspire DFX for metal part sorting.
Source raw episodes: episode_0311 through episode_0330 from local metal_part_sort_v10
Episodes: 20
Frames: 9372
Cameras: external, left wrist, right wrist, plus left_high/head slot retained for compatibility
GR00T training uses external as the logical head camera input
State/action: 26 dims (left arm 7, right arm 7, left hand 6, right hand 6)
pi-dev-plugins
Pi Coding Agent Plugins
Dataset contains metadata for approximately 5400 Pi Coding Agent plugins gathered from https://pi.dev/packages on 2026-09-21. Row example:
{
"name": "pi-mcp-adapter",
"description": "MCP (Model Context Protocol) adapter extension for Pi coding agent",
"types": [
"extension"
],
"author": "nicopreme",
"downloads_monthly": 972401,
"downloads_label": "972.4K/mo",
"published_label": "20m ago",
"published_ms": 1789973120934… See the full description on the dataset page: https://huggingface.co/datasets/kth8/pi-dev-plugins.frankapickplace_200episodesPidgin_ASR_Dataset_Combined
Naija-ASR-Corpus v1.0 (NAC-v1.0)
A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM)
📌 Dataset Summary
Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC).
The NAC Team processed the original long-form recordings by:
Segmenting the audio into sentence-level clips.
Transcribing/Aligning the text to create paired audio-text data suitable for ASR training.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/AnonXx/Pidgin_ASR_Dataset_Combined.stocks-PIDILITIND-1D-candlesdigitize-pid-symbols
Digitize-PID (Symbols only)
Note: I am not the author of this dataset
An annotated synthetic dataset, Dataset-P&ID, of 500 P&IDs with incorporates different
types of noise and complex symbols. This dataset contains only the symbols, i.e., under
the object detection task.
Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021). Digitize-PID: Automatic Digitization
of Piping and Instrumentation Diagrams. ArXiv, abs/2109.03794.
Original data repository:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-symbols.metal_part_sort_v5pidgin-asr-combined
Pidgin ASR Combined
A unified Nigerian Pidgin English speech-to-text dataset that combines
publicly available Pidgin ASR sources into a single train / validation /
test setup with a consistent schema. Built for fine-tuning Whisper-family
models on Nigerian Pidgin (Naija, pcm).
~8.6 hours, 4,278 clips, 10 source speakers, 16 kHz mono WAV.
Used to train michaelodafe/whisper-pidgin-v1
(21.37% WER on the test split, beating the published Wav2Vec2-XLSR-53
baseline by 8.2 pp).… See the full description on the dataset page: https://huggingface.co/datasets/michaelodafe/pidgin-asr-combined.
