datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.bengali-diarization-synthetic-v3synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.danish-diarization-bench
Danish Diarization Benchmark (Synthetic) — v2
A 3996-row synthetic speaker-diarization benchmark in Danish, built by mixing
single-speaker utterances from
syvai/danish-asr-unified
into multi-speaker recordings.
What changed in v2 (2026-05-18)
Per-segment text — each entry in segments now carries its text field directly. The redundant parallel texts column has been removed. Old consumers that joined segments[i] with texts[i] should switch to segments[i]["text"].
Silent… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-diarization-bench.coral-diarization-trainsada-diarization-preview
SADA 2022 Arabic Diarization
Training-ready speaker-attributed ASR windows derived from
SADA 2022. The source
recordings are mirrored at
khaledalganem/sada2022.
Splits
train: 36,004 windows, 202.064 hours, 4,062 recordings
validation: 853 windows, 4.774 hours, 88 recordings
test: 901 windows, 5.006 hours, 111 recordings
Total: 37,758 windows and
211.844 hours.
The official SADA train, validation, and test partitions are preserved.
Windows are 8–28 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Mohaddz/sada-diarization-preview.bengali-diarization-synthetic-v4
Bengali Speaker Diarization Synthetic Dataset V4
Synthetic Bengali speaker diarization dataset with natural overlapping speech patterns using timeline-based random chunk placement.
Dataset Overview
Property
Value
Total Samples
600
Speaker Categories
1-30 speakers per sample
Samples per Category
20
Duration per Sample
~30 minutes
Total Duration
~300 hours
Sample Rate
16000 Hz
Format
WAV (audio) + RTTM (labels) + JSON (metadata)… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-diarization-synthetic-v4.gdrive-sbpn-fresh-diarization-colab-l4-20260813-benchmarkgdrive-sbpn-fresh-diarization-optimized-terminal-l4-20260814-benchmarkgdrive-sbpn-fresh-diarization-h100-20260815-benchmarkCATS-ami-speaker-diarization
CATS-ami-speaker-diarization Dataset
Overview
This dataset is designed for speaker diarization tasks on the CATS (Comprehensive Assesment for Testing Speech) Dataset. It contains audio segments from the AMI Meeting Corpus with corresponding transcriptions and speaker information.
Dataset Structure
Each example in the dataset contains:
meeting_id: Identifier for the source meeting
label: Segment identifier within the meeting
start_time/end_time: Timestamp… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/CATS-ami-speaker-diarization.
