datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
French_MultiSpeaker_Diarization
French Multi-Speaker Diarization Dataset
This dataset is designed for training models for multi-speaker diarization in French. It contains transcriptions of conversations with multiple speakers, where the dialogues have been segmented and labeled by speaker. The dataset is ideal for tasks such as speaker diarization.
Note: All conversations in this dataset are entirely fictitious and were generated using AI. They do not reference real events, people, or organizations.… See the full description on the dataset page: https://huggingface.co/datasets/olafdil/French_MultiSpeaker_Diarization.gdrive-sbpn-fresh-diarization-colab-l4-20260813-benchmarkgdrive-sbpn-supplied-diarization-terminal-l4-20260814-benchmarkgdrive-sbpn-fresh-diarization-optimized-terminal-l4-20260814-benchmarkgdrive-sbpn-fresh-diarization-h100-20260815-benchmark
