CoolFace
Datasetpublic

yangwang825/vox2-veri-full

VoxCeleb 2 VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube. Verification Split train validation test # of speakers 5,994 5,994 118 # of samples 982,808 109,201 36,237 Data Fields ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>. duration (float64): The duration of the segment in seconds. wav (string): The filepath of the waveform. start… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-full.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes31downloads
Dataset Card

VoxCeleb 2

VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.

Verification Split

trainvalidationtest
# of speakers5,9945,994118
# of samples982,808109,20136,237

Data Fields

  • ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
  • duration (float64): The duration of the segment in seconds.
  • wav (string): The filepath of the waveform.
  • start (int64): The start index of the segment, which is (start seconds) × (sample rate).
  • stop (int64): The stop index of the segment, which is (stop seconds) × (sample rate).
  • spk_id (string): The ID of the speaker.

Example:

{
  'ID': 'id09056--00112_0_89088',
  'duration': 5.568,
  'wav': 'id09056/U2mRgZ1tW04/00112.wav',
  'start': 0,
  'stop': 89088,
  'spk_id': 'id09056'
}

References

  • https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox2.html