datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.power-plant-olmoearth-segmentation
OlmoEarth v1.2-ready power energy dataset
Derived cook 20260830T191500Z from pinned source 35da3d0549b16e15114ded1eb7df3f4a378e9b4f.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12453
Quality-unavailable rows: 1
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 10
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/power-plant-olmoearth-segmentation.droid_dataset_segmentation_mask
DROID SAM 3.1 Segmentation Masks
This dataset is a mask-only sidecar generated from the original
droid_101/0.0.1 RLDS release. It does not redistribute DROID images or
actions. Its episode_index follows the RLDS episode order.
The same episodes appear in lerobot/droid_1.0.1, but LeRobot stores them in a
different episode order. Therefore, mask episode_index and LeRobot
episode_index must not be joined directly. Use the mapping file described
below to associate these masks with… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/droid_dataset_segmentation_mask.table_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.recitation-segmentation-augmented
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
Paper | Project Page | Code
Introduction
This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated utterances).… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation-augmented.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation.im3-datacenter-olmoearth-segmentation
OlmoEarth v1.2-ready datacenter energy dataset
Derived cook 20260830T191500Z from pinned source 1f3ee5f334dd53940d088a89db57aed62cc925d5.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 1417
Quality-unavailable rows: 0
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 1
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3-datacenter-olmoearth-segmentation.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.recitation-segmentation-augmented
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
Paper | Project Page | Code
Introduction
This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation-augmented.wind-and-solar-candidate-olmoearth-segmentation
OlmoEarth v1.2-ready renewable energy dataset
Derived cook 20260830T191500Z from pinned source 2c551f58998cd25554ea679148b21a9c701b51db.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12649
Quality-unavailable rows: 4
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 4
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/wind-and-solar-candidate-olmoearth-segmentation.africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria
Customer Segmentation Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria.stereotactic-radiosurgery-k1-with-segmentation
🎯 Stereotactic Radiosurgery Dataset (SRS)
🏥 400 synthetic patient records describing the clinical, imaging, segmentation, and treatment-planning metadata of a stereotactic radiosurgery workflow, delivered as a single CSV with placeholder file paths.
⚠️ Disclaimer: This is a metadata-only synthetic dataset. It contains no real patients, no image files, and no segmentation files. Every record is generated; paths in the imaging and segmentation columns are placeholders that do… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/stereotactic-radiosurgery-k1-with-segmentation.image-segmentation-checkpoint-downloadsclinical-healing-trajectory-tokenization-phase-segmentation-v0.1What this dataset tests
Whether a model can segment high-frequency recovery datainto interpretable healing phases.
Required outputs
phase_sequence
phase_boundaries
phase_confidence_0_100
Token labels
acute_drop
early_rebound
consolidation_plateau
oscillatory_instability
secondary_drop
delayed_rebound
steady_ascent
maladaptive_plateau
recovery_lock_in
Boundary format
Use day indicesexampleacute_drop d0-d2
Typical failures
naming phases without boundaries… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-tokenization-phase-segmentation-v0.1.abisheksudarshan_customer-segmentation
Customer Segmentation
Multiclass Classification
Dataset Info
Source: Kaggle
Original Size: 0.10 MB
Kaggle Downloads: 9,253
Files: 2
Files
test.csv
train.csv
Mirrored from Kaggle
Image-to-Segmentation
Segmentation Data Subset
image_id: Refer to image_id on image_data subset
segmentation_id: Segmentation identifier
segmentation_information: COCO Format annotations: [[x1, y1, x2, y2, x3, y3, x4, y4]]
Field-Segmentationcmu_panoptic_dataset_instance_segmentation_mask
CMU Panoptic person instance masks (SAM 3.1 pseudo-labels)
TODO before making this repo public: confirm permission from the CMU Panoptic
Studio organisers. Until then the repo must stay gated and private. See
LICENSE_NOTICE.md.
Per-person instance segmentation masks for the CMU Panoptic Studio dataset, generated
with SAM 3.1 (facebook/sam3.1 @ daa63191845a41281374e725f4c9e51c7a824460) and aligned to the
dataset's own 3D skeletons: masks and pose share the same person id across… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/cmu_panoptic_dataset_instance_segmentation_mask.Segmentation-drivablecar-segmentation-assignmentThe paquet file is organized as a table. The columns read from left to right: Row number, frame index, timestamp, car component, confidenct score, and the two x-axis
and two y-axis where one of the corner of the boxes might reside. The paquet file used a Yolo8 object detector to find car components in various frames
in the given car video.
NALA-segmentation-ID
