DJRHails/pyannote-embedding-commonvoice-en
Pre-computed speaker embeddings Pre-computed 512-dim L2-normalized speaker embeddings extracted with pyannote/embedding over commonvoice-en. One utterance per speaker, minimum 3 s duration. Contents commonvoice-en.pyannote-embedding.npz — numpy .npz archive with: embeddings: (5000, 512) float32 speaker_ids: (5000,) string IDs from the source corpus metadata_json: per-speaker metadata (accent / age / gender / source URL) — populated for 5000 / 5000 speakers… See the full description on the dataset page: https://huggingface.co/datasets/DJRHails/pyannote-embedding-commonvoice-en.
Pre-computed speaker embeddings
Pre-computed 512-dim L2-normalized speaker embeddings extracted with `pyannote/embedding` over commonvoice-en. One utterance per speaker, minimum 3 s duration.
Contents
commonvoice-en.pyannote-embedding.npz— numpy.npzarchive with:embeddings:(5000, 512)float32speaker_ids:(5000,)string IDs from the source corpusmetadata_json: per-speaker metadata (accent / age / gender / source URL) — populated for 5000 / 5000 speakersn_speakers,sourcefor provenance
Loading
import numpy as np
data = np.load("commonvoice-en.pyannote-embedding.npz", allow_pickle=True)
embeddings = data["embeddings"] # (N, 512)
speaker_ids = list(data["speaker_ids"]) # length NRegenerating
This file was produced by `voxpath` via:
voxpath corpus build commonvoice --max-speakers 5000 \
--output commonvoice-en.pyannote-embedding.npzvoxpath corpus build streams the source audio, embeds each speaker's first valid (≥ 3 s) utterance with pyannote/embedding, L2-normalises, and writes the .npz.
Why model-specific
Speaker embeddings are not portable across embedders. A wespeaker embedding and a pyannote/embedding embedding for the same audio lie in different spaces and can't be compared or quantized together. This repo is named after the embedding model so users can find the right artifact for their pipeline at a glance.
