datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vhr-building-segmentation
HOT Building Segmentation Dataset
Dataset Description
A semantic segmentation dataset for building footprint extraction from aerial imagery, built from validated Humanitarian OpenStreetMap Team (HOT) Tasking Manager projects that use OpenAerialMap (OAM) imagery.
Dataset Summary
This dataset pairs 256x256 aerial image tiles (zoom level 19) from OpenAerialMap with building footprint labels from OpenStreetMap. All source projects have been fully… See the full description on the dataset page: https://huggingface.co/datasets/hotosm/vhr-building-segmentation.power-plant-olmoearth-segmentation
OlmoEarth v1.2-ready power energy dataset
Derived cook 20260830T191500Z from pinned source 35da3d0549b16e15114ded1eb7df3f4a378e9b4f.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12453
Quality-unavailable rows: 1
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 10
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/power-plant-olmoearth-segmentation.droid_dataset_segmentation_mask
DROID SAM 3.1 Segmentation Masks
This dataset is a mask-only sidecar generated from the original
droid_101/0.0.1 RLDS release. It does not redistribute DROID images or
actions. Its episode_index follows the RLDS episode order.
The same episodes appear in lerobot/droid_1.0.1, but LeRobot stores them in a
different episode order. Therefore, mask episode_index and LeRobot
episode_index must not be joined directly. Use the mapping file described
below to associate these masks with… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/droid_dataset_segmentation_mask.table_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.hot-building-segmentation
HOT Building Segmentation Dataset
Dataset Description
A semantic segmentation dataset for building footprint extraction from aerial imagery, built from validated Humanitarian OpenStreetMap Team (HOT) Tasking Manager projects that use OpenAerialMap (OAM) imagery.
Dataset Summary
This dataset pairs 256x256 aerial image tiles (zoom level 19) from OpenAerialMap with building footprint labels from OpenStreetMap. All source projects have been fully validated through… See the full description on the dataset page: https://huggingface.co/datasets/kshitijrajsharma/hot-building-segmentation.fashion_segmentationim3-datacenter-olmoearth-segmentation
OlmoEarth v1.2-ready datacenter energy dataset
Derived cook 20260830T191500Z from pinned source 1f3ee5f334dd53940d088a89db57aed62cc925d5.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 1417
Quality-unavailable rows: 0
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 1
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3-datacenter-olmoearth-segmentation.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation.genatator-gene-segmentation-dataset
Overview
genatator-gene-segmentation-dataset is a nucleotide-level gene segmentation dataset designed for training and evaluating DNA language models and related sequence models on transcript structure prediction. The dataset targets biologically detailed reconstruction of transcript architecture and supports benchmarking in the context of ab initio gene annotation. Each example represents one annotated transcript and provides nucleotide-resolution labels describing transcript… See the full description on the dataset page: https://huggingface.co/datasets/AIRI-Institute/genatator-gene-segmentation-dataset.wind-and-solar-candidate-olmoearth-segmentation
OlmoEarth v1.2-ready renewable energy dataset
Derived cook 20260830T191500Z from pinned source 2c551f58998cd25554ea679148b21a9c701b51db.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12649
Quality-unavailable rows: 4
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 4
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/wind-and-solar-candidate-olmoearth-segmentation.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation.brain-tumor-image-dataset-semantic-segmentation
Dataset Card for "brain-tumor-image-dataset-semantic-segmentation"
Dataset Description
The Brain Tumor Image Dataset (BTID) for Semantic Segmentation contains MRI images and annotations aimed at training and evaluating segmentation models. This dataset was sourced from Kaggle and includes detailed segmentation masks indicating the presence and boundaries of brain tumors.
This dataset can be used for developing and benchmarking algorithms for medical image segmentation… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/brain-tumor-image-dataset-semantic-segmentation.et_handwriting_segmentation
Dataset of Text Region and Line Coordinates in Handwritten Estonian Documents
Dataset Description
This dataset contains coordinate annotations for Estonian historical documents, extracted from Transkribus exports. It includes full page images with precise coordinate information for text regions and text lines, designed for text detection, layout analysis, and document structure understanding tasks.
📊 Dataset Summary
Total Examples: 7,664 images
Language: 🇪🇪… See the full description on the dataset page: https://huggingface.co/datasets/Rahvusarhiiv/et_handwriting_segmentation.streetview_segmentationswelding-component-segmentationpengwin-2026-anatomy-segmentation-pelviccoco2017-segmentation-50k-256x256
📄 License and Attribution
This dataset is a downsampled version of the COCO 2017 dataset, tailored for segmentation tasks. It has the following fields:
image: 256x256 image
segmentation: 256x256 image. Each pixel encodes the class of that pixel. See class_names_dict.json for a legend.
captions: a list of captions for the image, each by a different labeler.
Use the dataset as follows:
import requests
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/peteole/coco2017-segmentation-50k-256x256.pengwin-2026-anatomy-segmentation-pelvic-full-resgenerated-passports-segmentation
GENERATED USA Passports Segmentation
The dataset contains a collection of images representing GENERATED USA Passports. Each passport image is segmented into different zones, including the passport zone, photo, name, surname, date of birth, sex, nationality, passport number, and MRZ (Machine Readable Zone).
The dataset can be utilized for computer vision, object detection, data extraction and machine learning models.
Generated passports can assist in conducting research without… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/generated-passports-segmentation.satellite-building-segmentation
Dataset Labels
['building']
Number of Images
{'train': 6764, 'valid': 1934, 'test': 967}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("keremberke/satellite-building-segmentation", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/roboflow-universe-projects/buildings-instance-segmentation/dataset/1
Citation… See the full description on the dataset page: https://huggingface.co/datasets/merve/satellite-building-segmentation.weedy_rice_segmentation
Weedy Rice Segmentation
A dataset for semantic segmentation of weedy rice. The dataset contains 734 images with pixel-level mask annotations.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
The original train/test/val split has been preserved in the split column.
Citation
@article{nguyen2025dataset,
title={A dataset of aligned RGB and multispectral UAV imagery for semantic segmentation of weedy rice}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/weedy_rice_segmentation.pengwin-2026-anatomy-segmentation-femurSegmentation_Reward_DPO_condvineyard_grape_segmentation
Vineyard Grape Segmentation
A dataset for semantic segmentation of grape bunches in a vineyard. The dataset contains 646 images with pixel-level mask annotations.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{ariza2023dataset,
title={Dataset on UAV RGB videos acquired over a vineyard including bunch labels for object detection and tracking},
author={Ariza-Sent{\'\i}s, Mar and V{\'e}lez, Sergio… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/vineyard_grape_segmentation.elwha-segmentation-predictsentence-segmentation-dpo-raw
Dataset Card for "sentence-segmentation-dpo-raw"
More Information needed
strawberry_runner_segmentation
Strawberry Runner Segmentation
A dataset for semantic segmentation of strawberry runners. The dataset contains 3,345 images with pixel-level mask annotations.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
The original train/test/val split has been preserved in the split column.
Citation
@article{zhou2025deep,
title={Deep learning for strawberry runner detection integrating ground and aerial imaging}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/strawberry_runner_segmentation.Sesame_Aerial_weed_segmentation
Sesame Aerial Weed Segmentation
This dataset provides real aerial RGB imagery of tobacco and sesame crop fields in Pakistan, captured using a drone-mounted camera for weed segmentation tasks. The images were collected in field environments across multiple locations, offering diverse examples of weed distribution patterns in agricultural settings. The dataset contains 160 images with pixel-level mask annotations.
This dataset is indexed on https://project-agml.github.io/ as part… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Sesame_Aerial_weed_segmentation.pheno4d_point_cloud_segmentation
Pheno4D Point Cloud Segmentation
A dataset for point cloud semantic segmentation of maize and tomato plants. The dataset contains 126 labeled 3D scans across two species, tomato (single label scheme) and maize (dual label scheme: collar-based and tip-based organ boundaries), with per-point labels identifying plant organs (e.g. stem, leaf).
Each sample includes:
points: an (N, 3) array of x, y, z point coordinates, N varies per scan
mask: an (N, 1) array of per-point organ… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/pheno4d_point_cloud_segmentation.floorplan-room-segmentation
Floorplans Dataset
This dataset is derived from the Floorplans Diff dataset and has been curated by removing all unannotated images to ensure clean and consistent training data.
It is designed for semantic image segmentation, specifically focusing on identifying and segmenting rooms within floorplan images.
Each sample consists of an image paired with a corresponding segmentation mask, enabling models to learn pixel-level classification for the room class.
Origin
This… See the full description on the dataset page: https://huggingface.co/datasets/peaceAsh/floorplan-room-segmentation.
