datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-speaker-diarization-dataset-fa-large-3000synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.mac-m4pro-fresh-diarization-demucs-20260902
gdrive-sbpn-fresh-diarization-demucs-mac-m4pro-20260902
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/mac-m4pro-fresh-diarization-demucs-20260902.danish-diarization-bench
Danish Diarization Benchmark (Synthetic) — v2
A 3996-row synthetic speaker-diarization benchmark in Danish, built by mixing
single-speaker utterances from
syvai/danish-asr-unified
into multi-speaker recordings.
What changed in v2 (2026-05-18)
Per-segment text — each entry in segments now carries its text field directly. The redundant parallel texts column has been removed. Old consumers that joined segments[i] with texts[i] should switch to segments[i]["text"].
Silent… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-diarization-bench.sada-diarization-preview
SADA 2022 Arabic Diarization
Training-ready speaker-attributed ASR windows derived from
SADA 2022. The source
recordings are mirrored at
khaledalganem/sada2022.
Splits
train: 36,004 windows, 202.064 hours, 4,062 recordings
validation: 853 windows, 4.774 hours, 88 recordings
test: 901 windows, 5.006 hours, 111 recordings
Total: 37,758 windows and
211.844 hours.
The official SADA train, validation, and test partitions are preserved.
Windows are 8–28 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Mohaddz/sada-diarization-preview.gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814
gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814.gdrive-sbpn-supplied-diarization-demucs-terminal-l4-20260814
gdrive-sbpn-supplied-diarization-demucs-terminal-l4-20260814
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-supplied-diarization-demucs-terminal-l4-20260814.gdrive-sbpn-fresh-diarization-demucs-h100-20260815
gdrive-sbpn-fresh-diarization-demucs-h100-20260815
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-demucs-h100-20260815.gdrive-sbpn-diarization-before-after-review-20260812
gdrive-sbpn-diarization-before-after-review-20260812
Each row is one transcript row whose timestamp changed during diarization boundary correction. Play audio_before_diarization and audio_after_diarization side by side. Both are cut from the same original recording; Demucs is not used. Text and speaker are unchanged.
start_movement_seconds and end_movement_seconds equal after minus before: negative means earlier and positive means later. boundary_audit_json records the fusion… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-diarization-before-after-review-20260812.
