x-vector
Datasets
All datasets matching “x-vector”cmu-arctic-xvectors
Speaker embeddings extracted from CMU ARCTIC
There is one .npy file for each utterance in the dataset, 7931 files in total. The speaker embeddings are 512-element X-vectors.
The CMU ARCTIC dataset divides the utterances among the following speakers:
bdl (US male)
slt (US female)
jmk (Canadian male)
awb (Scottish male)
rms (US male)
clb (US female)
ksp (Indian male)
The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model.
Usage:
from… See the full description on the dataset page: https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors.cmu-arctic-xvectors-extractedarabic_xvector_embeddings
Arabic Speaker Embeddings extracted from ASC and ClArTTS
There is one speaker embedding for each utterance in the validation set of both datasets. The speaker embeddings are 512-element X-vectors.
Arabic Speech Corpus has 100 files for a single male speaker and ClArTTS has 205 files for a single male speaker.
The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model.
Usage:
from datasets import load_dataset
embeddings_dataset =… See the full description on the dataset page: https://huggingface.co/datasets/herwoww/arabic_xvector_embeddings.cmu-arctic-xvectorsThe CMU ARCTIC dataset without needing to run remote code, so it is compatible with datasets >= 4.0.0.
xvector_nnet_1a_libritts_clean_460This dataset contains x-vectors extracted using kaldi toolkit from libritts-{clean-460,dev-clean,test-clean} using the pre-trained model from http://kaldi-asr.org/models/8/0008_sitw_v2_1a.tar.gz
sa-xvector
