datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
p1-segments
DR P1 speech segments
Dataset
Danish speech clips from DR P1, in mono 16 kHz OGG/Opus, with verbatim text, timing, and speaker metadata. Transcript text and speaker attribution may contain automated errors.
Source
The recordings cover roughly 2006–2022 and come from DR P1 recordings in kb.dk’s DR archive. Audio is sourced through the pinned syvai/p1 revision 449b9c2294026df6d0d37538f279fdec03f565ff. Transcripts were generated with ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1-segments.warsh-segments-v3
Haitam03/warsh-v3
Warsh (Rewayat Warsh A'n Nafi') Quran recitation, segmented at waqf with
obadx/recitation-segmenter-v2.
Built with warsh-data.
Layout
path
what
data/<reciter>/<surah>.parquet
one file per source recording, audio embedded as 16 kHz mono FLAC
raw/<reciter>/<surah>.mp3
the source recording it came from
segment_params.json
the settings this corpus was produced with
One parquet per source recording, named after it, so re-running a… See the full description on the dataset page: https://huggingface.co/datasets/Haitam03/warsh-segments-v3.cv-v1.0-segment
CommonVoice v1 Phone-Segment Alignments
Phone-level time alignments for 10 languages of Mozilla Common Voice,
packaged in a canonical segmentation schema with embedded 16 kHz audio. The
phone boundaries come from the charsiu/cv_ali
release of MFA alignments; the audio and transcripts come from
Common Voice Corpus 13.0 (2023-03-09).
Dataset summary
lang
train rows
train hrs
val rows
val hrs
test rows
test hrs
en
1,008,669
1,354.0
3,537
4.9
1,285
1.7
rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.ace-opencpop-segments
Citation Information
@misc{shi2024singingvoicedatascalingup,
title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing},
author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe},
year={2024},
eprint={2401.17619},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2401.17619},
}
news-segmentationlibrispeech-segment
LibriSpeech Segment
English read-speech corpus with phone-level time alignments (Montreal
Forced Aligner). Suitable for training and evaluating phone recognition and
phonetic segmentation models.
Sources
Audio: LibriSpeech (OpenSLR 12) by
Vassil Panayotov, Guoguo Chen, Daniel Povey, Sanjeev Khudanpur (2015).
Phone alignments:
anyspeech/librispeech_MFA_alignments.
Splits
Split
Utterances
train.clean.100
28,538
train.clean.360
104,008… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/librispeech-segment.ace-kising-segments
Citation Information
@misc{shi2024singingvoicedatascalingup,
title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing},
author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe},
year={2024},
eprint={2401.17619},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2401.17619},
}
parczech4speech-segmented
ParCzech4Speech (Sentence-Segmented Variant)
Dataset Summary
ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts.
This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries.
It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection.
Using WhisperX and Wav2Vec 2.0… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.musdb_segmentslrclib_segmentedrecitation-segmentation-augmented
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
Paper | Project Page | Code
Introduction
This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated utterances).… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation-augmented.vocal-bursts-segments
vocal-bursts-segments
128,165 vocal-burst segments from two acoustically unrelated sources, cut with one policy, plus
46,494 verified no-burst segments. Every segment is a single burst — a laugh, a
sigh, a gasp — and nothing else.
subtree
what it is
segments
real/
real speech (Emilia, LAION voice profiles, vocal-bursts-clean, Kartoffelphon), the 5,161 segments of laion/vocal-bursts-gemini-segments
5,161
dramabox/
synthetic voice-acting output, cut out of… See the full description on the dataset page: https://huggingface.co/datasets/laion/vocal-bursts-segments.SBC_segmented
Dataset Card for "SBC_segmented"
More Information needed
recitation-segmentation-augmented
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
Paper | Project Page | Code
Introduction
This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation-augmented.thchs30-segment
THCHS-30 Segment
Mandarin Chinese read-speech corpus with phone-level time alignments.
Suitable for training and evaluating phone recognition and phonetic
segmentation models.
Sources
Audio: THCHS-30 (OpenSLR 18) by
Dong Wang, Xuewei Zhang, Zhiyong Zhang (Tsinghua University, 2015).
Phone alignments:
anyspeech/THCHS-30-alignments.
Splits
Split
Utterances
train
10,000
val
893
test
2,495
Splits follow the original OpenSLR 18 directory… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/thchs30-segment.quran_segmentsvocal-bursts-gemini-segments
burst_gemini_segments (Dataset B)
5,161 vocal-burst segments cut out of 3,304 real speech utterances, one per
event that Gemini 3.8 Flash asserted. Each file is a single burst — a laugh, a sigh, a gasp — and
nothing else. Median length 0.76 s; 1.39 hours in total.
Built to retrain a burst classifier. The detector this project shipped emitted Shriek zero
times over a 60-clip audit, used 8 of its 83 labels, and put the requested burst in its top-3 on
3 of 60 clips. Every burst… See the full description on the dataset page: https://huggingface.co/datasets/laion/vocal-bursts-gemini-segments.warsh-segmentsLHCP-ASR-segments
LHCP-ASR Segments
This dataset is a segment-level distribution derived from mllp/LHCP-ASR (and the original LHCP-ASR repository), an English speech corpus for narrow-domain ASR benchmarking in high-energy particle physics.
Unlike previous versions, this repository provides audio directly at the segment level (<30 seconds each) for evaluation, adds talk-level metadata and cleans up transcription tags.
Differences from mllp/LHCP-ASR
No subsets: Directly formatted… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR-segments.spc_r_segmented
i4ds/spc_r_segmented
Diarized and segmented speech dataset derived from i4ds/spc_r.
Description
Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline:
Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap.
Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap between… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/spc_r_segmented.thomcles-persian-farsi-speech-whisper-segmented-under30sspc_r_segmented
eko57/spc_r_segmented
Diarized and segmented speech dataset derived from i4ds/spc_r.
Description
Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline:
Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap.
Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap… See the full description on the dataset page: https://huggingface.co/datasets/eko57/spc_r_segmented.LibriConvo-segmented
🗣️ LibriConvo-Segmented
LibriConvo-Segmented is a segmented version of the LibriConvo corpus — a simulated two-speaker conversational dataset built using Speaker-Aware Conversation Simulation (SASC).It is designed for training and evaluation of multi-speaker speech processing systems, including speaker diarization, automatic speech recognition (ASR), and overlapping speech modeling.
This segmented version provides ≤30-second conversational fragments derived from full LibriConvo… See the full description on the dataset page: https://huggingface.co/datasets/gedeonmate/LibriConvo-segmented.DEAM_stripped_vocals_segmentedspeech_segmentslibrispeech-asr-whisper-segmented-under30snorthtts-men-v3-whisper-segmented-under30sfleurs-farsi-whisper-segmented-under30smedical-segmentation-dataset_v2modified-shemo-whisper-segmented-under30s
