datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
segments
NuBerea/segments
Analytical unit definitions for biblical, Second Temple, rabbinic, and early Christian
corpora: pericope boundaries for the Hebrew Bible and Greek New Testament, segment
boundaries for the Dead Sea Scrolls, Talmudic literature (Mishnah, Tosefta, Bavli,
Yerushalmi), early Christian writings, Nag Hammadi codices, Old Testament pseudepigrapha,
and Migne's Patrologia Latina, together with a cross-corpus event topology (canonical
biblical events, their aliases… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/segments.power-plant-olmoearth-segmentation
OlmoEarth v1.2-ready power energy dataset
Derived cook 20260830T191500Z from pinned source 35da3d0549b16e15114ded1eb7df3f4a378e9b4f.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12453
Quality-unavailable rows: 1
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 10
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/power-plant-olmoearth-segmentation.droid_dataset_segmentation_mask
DROID SAM 3.1 Segmentation Masks
This dataset is a mask-only sidecar generated from the original
droid_101/0.0.1 RLDS release. It does not redistribute DROID images or
actions. Its episode_index follows the RLDS episode order.
The same episodes appear in lerobot/droid_1.0.1, but LeRobot stores them in a
different episode order. Therefore, mask episode_index and LeRobot
episode_index must not be joined directly. Use the mapping file described
below to associate these masks with… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/droid_dataset_segmentation_mask.phage-segmentdbsit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199table_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.im3-datacenter-olmoearth-segmentation
OlmoEarth v1.2-ready datacenter energy dataset
Derived cook 20260830T191500Z from pinned source 1f3ee5f334dd53940d088a89db57aed62cc925d5.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 1417
Quality-unavailable rows: 0
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 1
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3-datacenter-olmoearth-segmentation.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation.wind-and-solar-candidate-olmoearth-segmentation
OlmoEarth v1.2-ready renewable energy dataset
Derived cook 20260830T191500Z from pinned source 2c551f58998cd25554ea679148b21a9c701b51db.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12649
Quality-unavailable rows: 4
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 4
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/wind-and-solar-candidate-olmoearth-segmentation.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation.sit-latents-ode-heun-1000-class-0_1000-samples-segment-400-499Segmentation_de_la_feuille_de_papier_qui_sert_de_zone_20260924_110415This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Tridex/Segmentation_de_la_feuille_de_papier_qui_sert_de_zone_20260924_110415.africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria
Customer Segmentation Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria.tech-talks-segments
Awesome Tech Talks Dataset
A curated dataset of 2,200+ technical sessions, workshops, and keynotes from official engineering organizations including Google, Microsoft, OpenAI, Anthropic, and Cursor. The dataset includes video metadata, structured topic classifications, cleaned transcripts, and 42,000+ segmented text chunks designed for Retrieval-Augmented Generation (RAG), vector search, and language model evaluation.
Dataset Summary
Attribute
Value… See the full description on the dataset page: https://huggingface.co/datasets/0xShashi/tech-talks-segments.sit-latents-ode-heun-segment-0-99kising_score_segmentssit-latents-ode-heun-1000-class-0_1000-samples-segment-500-599sit-latents-ode-heun-1000-class-0_1000-samples-segment-600-699PhaStyle-SegmentDBimage-segmentation-checkpoint-downloadssit-latents-ode-heun-1000-test2-class-0_1000-samples-segment-10-19sit-latents-ode-heun-1000-class-0_1000-samples-segment-700-799sit-latents-ode-heun-1000-class-0_1000-samples-segment-900-999oc_segmentThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 22449,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RaspberryVitriol/oc_segment.Segmented_slightly_clear_subsetsit-latents-ode-heun-1000-test2-class-0_1000-samples-segment-20-29sit-latents-ode-heun-1000-class-0_1000-samples-segment-800-899Image-to-Segmentation
Segmentation Data Subset
image_id: Refer to image_id on image_data subset
segmentation_id: Segmentation identifier
segmentation_information: COCO Format annotations: [[x1, y1, x2, y2, x3, y3, x4, y4]]
sit-latents-ode-heun-1000-class-0_1000-samples-segment-200-299sit-latents-ode-heun-1000-class-0_1000-samples-segment-300-399
