pyannote
Datasets
All datasets matching “pyannote”pyannote-hindi-diarizationpyannote-embedding-librispeech-multi
Pre-computed speaker embeddings
Pre-computed 512-dim L2-normalized speaker embeddings extracted with
pyannote/embedding (512-dim) over
LibriSpeech train-clean-100 (all 251 speakers, 10 utterances each). 2510 utterances across 251 speakers, minimum 3 s duration.
Contents
librispeech-multi.pyannote-embedding.npz — numpy .npz archive with:
embeddings: (2510, 512) float32
speaker_ids: (2510,) string IDs from the source corpus
metadata_json: per-speaker metadata… See the full description on the dataset page: https://huggingface.co/datasets/DJRHails/pyannote-embedding-librispeech-multi.pyannote-embedding-voxceleb
Pre-computed speaker embeddings
Pre-computed 512-dim L2-normalized speaker embeddings extracted with
pyannote/embedding over
VoxCeleb 2 dev (5800 speakers via gaunernst/voxceleb2-dev-wds). One utterance per speaker, minimum 3 s duration.
Contents
voxceleb.pyannote-embedding.npz — numpy .npz archive with:
embeddings: (5800, 512) float32
speaker_ids: (5800,) string IDs from the source corpus
metadata_json: per-speaker metadata (accent / age / gender / source URL)
—… See the full description on the dataset page: https://huggingface.co/datasets/DJRHails/pyannote-embedding-voxceleb.pyannote-embedding-librispeech
Pre-computed speaker embeddings
Pre-computed 512-dim L2-normalized speaker embeddings extracted with
pyannote/embedding over
LibriSpeech train.100 + train.360 (1172 speakers via openslr/librispeech_asr). One utterance per speaker, minimum 3 s duration.
Contents
librispeech.pyannote-embedding.npz — numpy .npz archive with:
embeddings: (3507, 512) float32
speaker_ids: (3507,) string IDs from the source corpus
metadata_json: per-speaker metadata (accent / age / gender /… See the full description on the dataset page: https://huggingface.co/datasets/DJRHails/pyannote-embedding-librispeech.pyannote-embedding-commonvoice-en
Pre-computed speaker embeddings
Pre-computed 512-dim L2-normalized speaker embeddings extracted with
pyannote/embedding over
commonvoice-en. One utterance per speaker, minimum 3 s duration.
Contents
commonvoice-en.pyannote-embedding.npz — numpy .npz archive with:
embeddings: (5000, 512) float32
speaker_ids: (5000,) string IDs from the source corpus
metadata_json: per-speaker metadata (accent / age / gender / source URL)
— populated for 5000 / 5000 speakers
n_speakers… See the full description on the dataset page: https://huggingface.co/datasets/DJRHails/pyannote-embedding-commonvoice-en.pyannoteAI-EMMA-20k
pyannoteAI-EMMA-20k dataset
In the framework of the EMMA project, part of JSALT 2025, the pyannoteAI research lab ran its internal speaker diarization pipeline on the whole YODAS2 dataset.
From the output, we then selected a subset of the predictions for a total amount of around 20k hours of audio.
Licence
This dataset is licensed under CC BY-NC-SA 4.0 and therefore does not allow commercial use (e.g. training a commercial model using this dataset).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hbredin/pyannoteAI-EMMA-20k.
