datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
p1-segments
DR P1 speech segments
Dataset
Danish speech clips from DR P1, in mono 16 kHz OGG/Opus, with verbatim text, timing, and speaker metadata. Transcript text and speaker attribution may contain automated errors.
Source
The recordings cover roughly 2006–2022 and come from DR P1 recordings in kb.dk’s DR archive. Audio is sourced through the pinned syvai/p1 revision 449b9c2294026df6d0d37538f279fdec03f565ff. Transcripts were generated with ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1-segments.bert-base-uncased-refined-web-segment0
Dataset Card for "bert-base-uncased-refined-web-segment0"
More Information needed
warsh-segments-v3
Haitam03/warsh-v3
Warsh (Rewayat Warsh A'n Nafi') Quran recitation, segmented at waqf with
obadx/recitation-segmenter-v2.
Built with warsh-data.
Layout
path
what
data/<reciter>/<surah>.parquet
one file per source recording, audio embedded as 16 kHz mono FLAC
raw/<reciter>/<surah>.mp3
the source recording it came from
segment_params.json
the settings this corpus was produced with
One parquet per source recording, named after it, so re-running a… See the full description on the dataset page: https://huggingface.co/datasets/Haitam03/warsh-segments-v3.cv-v1.0-segment
CommonVoice v1 Phone-Segment Alignments
Phone-level time alignments for 10 languages of Mozilla Common Voice,
packaged in a canonical segmentation schema with embedded 16 kHz audio. The
phone boundaries come from the charsiu/cv_ali
release of MFA alignments; the audio and transcripts come from
Common Voice Corpus 13.0 (2023-03-09).
Dataset summary
lang
train rows
train hrs
val rows
val hrs
test rows
test hrs
en
1,008,669
1,354.0
3,537
4.9
1,285
1.7
rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.tmmluplus_CKIP_segmentedobject-segmentationace-opencpop-segments
Citation Information
@misc{shi2024singingvoicedatascalingup,
title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing},
author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe},
year={2024},
eprint={2401.17619},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2401.17619},
}
vhr-building-segmentation
HOT Building Segmentation Dataset
Dataset Description
A semantic segmentation dataset for building footprint extraction from aerial imagery, built from validated Humanitarian OpenStreetMap Team (HOT) Tasking Manager projects that use OpenAerialMap (OAM) imagery.
Dataset Summary
This dataset pairs 256x256 aerial image tiles (zoom level 19) from OpenAerialMap with building footprint labels from OpenStreetMap. All source projects have been fully… See the full description on the dataset page: https://huggingface.co/datasets/hotosm/vhr-building-segmentation.segments
NuBerea/segments
Analytical unit definitions for biblical, Second Temple, rabbinic, and early Christian
corpora: pericope boundaries for the Hebrew Bible and Greek New Testament, segment
boundaries for the Dead Sea Scrolls, Talmudic literature (Mishnah, Tosefta, Bavli,
Yerushalmi), early Christian writings, Nag Hammadi codices, Old Testament pseudepigrapha,
and Migne's Patrologia Latina, together with a cross-corpus event topology (canonical
biblical events, their aliases… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/segments.power-plant-olmoearth-segmentation
OlmoEarth v1.2-ready power energy dataset
Derived cook 20260830T191500Z from pinned source 35da3d0549b16e15114ded1eb7df3f4a378e9b4f.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12453
Quality-unavailable rows: 1
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 10
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/power-plant-olmoearth-segmentation.droid_dataset_segmentation_mask
DROID SAM 3.1 Segmentation Masks
This dataset is a mask-only sidecar generated from the original
droid_101/0.0.1 RLDS release. It does not redistribute DROID images or
actions. Its episode_index follows the RLDS episode order.
The same episodes appear in lerobot/droid_1.0.1, but LeRobot stores them in a
different episode order. Therefore, mask episode_index and LeRobot
episode_index must not be joined directly. Use the mapping file described
below to associate these masks with… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/droid_dataset_segmentation_mask.cityscapes_segmentationlibrispeech-segment
LibriSpeech Segment
English read-speech corpus with phone-level time alignments (Montreal
Forced Aligner). Suitable for training and evaluating phone recognition and
phonetic segmentation models.
Sources
Audio: LibriSpeech (OpenSLR 12) by
Vassil Panayotov, Guoguo Chen, Daniel Povey, Sanjeev Khudanpur (2015).
Phone alignments:
anyspeech/librispeech_MFA_alignments.
Splits
Split
Utterances
train.clean.100
28,538
train.clean.360
104,008… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/librispeech-segment.semantic-segmentation-test-sampleThis dataset contains 10 examples of the segments/sidewalk-semantic dataset (i.e. 10 images with corresponding ground-truth segmentation maps).
phage-segmentdblibero-pickandplace-segment-next-scene-ab-2ace-kising-segments
Citation Information
@misc{shi2024singingvoicedatascalingup,
title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing},
author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe},
year={2024},
eprint={2401.17619},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2401.17619},
}
SoccerNet_Field_SegmentationProcessed data from the Soccernet 2023 dataset. Processing notebook is included in this repo.
To see an example:
def show_item(item):
fig, axs = plt.subplots(nrows = 1, ncols = 4, figsize = (20, 4))
axs[0].imshow(item['image'])
axs[0].set_title("Image")
axs[0].axis('off')
axs[1].imshow(overlay_mask(item['image'], item['outlines']))
axs[1].set_title("Outlines")
axs[1].axis('off')
axs[2].imshow(show_segments(item['segments']))
axs[2].set_title("Segments")… See the full description on the dataset page: https://huggingface.co/datasets/nreHieW/SoccerNet_Field_Segmentation.libero-pickandplace-segmentThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1693,
"total_frames": 273465,
"total_tasks": 40,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10.0,
"splits": {
"train": "0:1693"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/libero-pickandplace-segment.uvh-coco-segmentedsit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199musdb_segmentslrclib_segmentedtable_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.agri-vision-2021-segmentationhot-building-segmentation
HOT Building Segmentation Dataset
Dataset Description
A semantic segmentation dataset for building footprint extraction from aerial imagery, built from validated Humanitarian OpenStreetMap Team (HOT) Tasking Manager projects that use OpenAerialMap (OAM) imagery.
Dataset Summary
This dataset pairs 256x256 aerial image tiles (zoom level 19) from OpenAerialMap with building footprint labels from OpenStreetMap. All source projects have been fully validated through… See the full description on the dataset page: https://huggingface.co/datasets/kshitijrajsharma/hot-building-segmentation.fashion_segmentationiMaterialist-2020-fashion-clothes-segmentation-train-part1tongue-images-384-segmented-augmentedim3-datacenter-olmoearth-segmentation
OlmoEarth v1.2-ready datacenter energy dataset
Derived cook 20260830T191500Z from pinned source 1f3ee5f334dd53940d088a89db57aed62cc925d5.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 1417
Quality-unavailable rows: 0
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 1
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3-datacenter-olmoearth-segmentation.
