datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.bengali-diarization-synthetic-v3echo-synthetic-diarization
Echo (Synthetic Set for Diarization)
This dataset is a synthetic dataset generated for evaluation of speaker
diarization models. It contains approximately two hours of speech data, each
file 60 seconds long, with and without overlap, with 2--5 speakers per file.
This dataset has been built using Echo.
luganda_callhome_diarization_dataset_MHDPvoicevox-diarization-ja
VOICEVOX合成 多話者日本語音声(話者分離検証用データセット)
Qwen3-ASR の話者分離ファインチューニング検証、特に embedding層+projectorのみの学習で話者分離が学べるか を確かめるために作成した、正解ラベルが厳密な合成データです。
作り方
gemma(gemma-4-31B-it)で話題別のモノローグ日本語テキストを生成
文単位に分割し、各文境界で確率0.3で話者を切替(2〜3話者)
各文を割当てた VOICEVOX 話者で音声合成し連結
→ 合成由来なので、各文の [開始秒, 終了秒] と話者が完全に既知(人手や別ASRに依存しない厳密ラベル)。
規模
698クリップ、各 約2分、opus (24kbps) / 16kHz 相当 mono
2話者 約88% / 3話者 約12%
話者は各クリップ内の登場順に spk_0, spk_1, spk_2 と採番
列 (data.jsonl)
列
内容… See the full description on the dataset page: https://huggingface.co/datasets/okadahiroaki/voicevox-diarization-ja.diarization_datasetshemo_diarization_datasetsynthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.synthetic-speaker-diarization-dataset-fa-large-3000sortformer-diarization-test-set
Sortformer Diarization Test Set
100 real speech samples extracted from LibriSpeech test-clean for speaker diarization testing and benchmarking with NVIDIA Sortformer 4spk-v2 ONNX models.
Usage with Sortformer ONNX
from huggingface_hub import snapshot_download
import soundfile as sf
# Download the test set
dataset_path = snapshot_download("DimQ1/sortformer-diarization-test-set")
# Load audio
audio, sr = sf.read(f"{dataset_path}/audio/ls_real_000.wav")… See the full description on the dataset page: https://huggingface.co/datasets/DimQ1/sortformer-diarization-test-set.speaker-diarization-rawsynthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.tamil-english-podcast-diarization
Tamil-English Code-Mixed Podcast Diarization Dataset
Dataset Summary
This dataset contains long-form Tamil-English code-mixed podcast recordings
annotated for speaker diarization research. The recordings consist of natural
conversational speech with multiple speakers and realistic acoustic conditions,
making the dataset suitable for evaluating diarization pipelines in
real-world scenarios.
The dataset is intended to support research in:
Speaker diarization
Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Rangasuthan/tamil-english-podcast-diarization.Speaker-Diarization-Instructions
Speaker-Diarization-Instructions
Convert diarization dataset from https://huggingface.co/diarizers-community into speech instructions dataset and chunk max to 30 seconds because most of speech encoder use for LLM come from Whisper Encoder.
We highly recommend to not include AMI test set from both AMI-IHM and AMI-SDM in training set to prevent contamination. This dataset supposely to become a speech diarization benchmark.
how to prepare the dataset
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Speaker-Diarization-Instructions.bangla-diarization-dataset-2librispeech-synthetic-speaker-diarization-datasetsynthetic-speaker-diarization-dataset-hindipyannote-hindi-diarizationsynthetic-speaker-diarization-datasetcallhome-eng-diarizationsynthetic-speaker-diarization-dataset-hindi-largedanish-diarization-bench
Danish Diarization Benchmark (Synthetic) — v2
A 3996-row synthetic speaker-diarization benchmark in Danish, built by mixing
single-speaker utterances from
syvai/danish-asr-unified
into multi-speaker recordings.
What changed in v2 (2026-05-18)
Per-segment text — each entry in segments now carries its text field directly. The redundant parallel texts column has been removed. Old consumers that joined segments[i] with texts[i] should switch to segments[i]["text"].
Silent… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-diarization-bench.ENNI_speaker_diarizationsynthetic-speaker-diarization-datasetcommon_voice_diarizationmac-m4pro-fresh-diarization-demucs-20260902
gdrive-sbpn-fresh-diarization-demucs-mac-m4pro-20260902
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/mac-m4pro-fresh-diarization-demucs-20260902.synthetic-speaker-diarization-dataset-fasynthetic dataset generated from persian common voice.
synthetic-diarization-italian-augmentedcoral-diarization-trainfandub-studio-diarization-ab-test
fandub-studio diarization A/B test
Test harness comparing two speaker-diarization approaches on a synthetic "fandub-like" clip (three voices + a music bed at -13 dB + noise at -30 dB), run 2026-09-20 on 2 CPU cores.
Result
Duration-weighted speaker attribution score (how much of the 66.2s of speech was assigned to the correct speaker, judged against ground truth):
pipeline
segments
score
A — repo spectral_vad algorithm (energy VAD + amplitude/ZCR… See the full description on the dataset page: https://huggingface.co/datasets/NathanG1en/fandub-studio-diarization-ab-test.
