datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vox1-veri-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
133777
14865
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
vox1-iden-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Identification Split
train
validation
test
# of speakers
1251
1251
1251
# of samples
138361
6904
8251
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
vox2-veri-full
VoxCeleb 2
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
5,994
5,994
118
# of samples
982,808
109,201
36,237
Data Fields
ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
duration (float64): The duration of the segment in seconds.
wav (string): The filepath of the waveform.
start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-full.
