datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.synthetic-speaker-diarization-dataset-fa-large-3000speaker-diarization-rawSpeaker-Diarization-Instructions
Speaker-Diarization-Instructions
Convert diarization dataset from https://huggingface.co/diarizers-community into speech instructions dataset and chunk max to 30 seconds because most of speech encoder use for LLM come from Whisper Encoder.
We highly recommend to not include AMI test set from both AMI-IHM and AMI-SDM in training set to prevent contamination. This dataset supposely to become a speech diarization benchmark.
how to prepare the dataset
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Speaker-Diarization-Instructions.librispeech-synthetic-speaker-diarization-datasetsynthetic-speaker-diarization-dataset-hindisynthetic-speaker-diarization-datasetsynthetic-speaker-diarization-dataset-hindi-largeENNI_speaker_diarizationsynthetic-speaker-diarization-datasetsynthetic-speaker-diarization-dataset-fasynthetic dataset generated from persian common voice.
CATS-ami-speaker-diarization-audiosynthetic-speaker-diarization-dataset-uzsynthetic-speaker-diarization-dataset-idsynthetic-speaker-diarization-dataset-hindisynthetic-speaker-diarization-dataset-nlSpeakerDiarization_SparseLibriMix2synthetic-ko-speaker-diarization-datasetsynthetic-speaker-diarization-dataset-hindi-shortBengali_Speaker_Diarization_Datasettemp-speaker-diarization-synthetic-datasetunsupervised-malay-youtube-speaker-diarization
Unsupervised malay speakers from youtube videos
10492 unique speakers with at least 75 hours of voice activities. Steps to reproduce at https://github.com/huseinzol05/malaya-speech/blob/master/data/youtube/process-youtube.ipynb
how-to
Download and extract processed-youtube.tar.gz, each processed videos saved as pickle, {video_name}.pkl.
Each pickle file got,
[{'wav_data':… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/unsupervised-malay-youtube-speaker-diarization.synthetic-speaker-diarization-dataset-hispeaker-diarizationspeaker-diarization-dataset-ami_corpus-sula-callhome-voxconverseBengali-Speaker-Diarization_datasetData Description
This is a speaker diarization dataset. The goal is to predict who spoke when in each audio file. Training data includes audio plus time-aligned speaker labels; test data includes audio only. Submissions list diarization segments for each test file.
Files
train/audio/.wav - training audio recordings
train/annotation/.csv - labels for each training audio (same base filename)
test/audio/.wav - test audio recordings to predict
sample_submission.csv - example submission in the… See the full description on the dataset page: https://huggingface.co/datasets/AdilShamim8/Bengali-Speaker-Diarization_dataset.synthetic-speaker-diarization-dataset-hindisynthetic-speaker-diarization-dataset-itahcf_speaker_diarization_datasetCATS-ami-speaker-diarization
CATS-ami-speaker-diarization Dataset
Overview
This dataset is designed for speaker diarization tasks on the CATS (Comprehensive Assesment for Testing Speech) Dataset. It contains audio segments from the AMI Meeting Corpus with corresponding transcriptions and speaker information.
Dataset Structure
Each example in the dataset contains:
meeting_id: Identifier for the source meeting
label: Segment identifier within the meeting
start_time/end_time: Timestamp… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/CATS-ami-speaker-diarization.
