speaker-diarization
synthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.synthetic-speaker-diarization-dataset-fa-large-3000speaker-diarization-rawSpeaker-Diarization-Instructions
Speaker-Diarization-Instructions
Convert diarization dataset from https://huggingface.co/diarizers-community into speech instructions dataset and chunk max to 30 seconds because most of speech encoder use for LLM come from Whisper Encoder.
We highly recommend to not include AMI test set from both AMI-IHM and AMI-SDM in training set to prevent contamination. This dataset supposely to become a speech diarization benchmark.
how to prepare the dataset
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Speaker-Diarization-Instructions.librispeech-synthetic-speaker-diarization-datasetsynthetic-speaker-diarization-dataset-hindi
