datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vox1-veri-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
133777
14865
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
vox1-iden-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Identification Split
train
validation
test
# of speakers
1251
1251
1251
# of samples
138361
6904
8251
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
vox1-iden-3s
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Identification Split
train
validation
test
# of speakers
1251
1251
1251
# of samples
306208
14479
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
Yemeni-Speech-Emotion-Dataset
YSED — Yemeni Speech Emotion Dataset (audio-classification repackaging)
A clean repackaging of YSED with a metadata.csv and stratified train/validation/test splits, for emotion classification on Yemeni Arabic.
Original dataset: Derhem, S., AL-Mekhlafi, E., AL-Majmar, N. A., & AL-Makhlafi, M. (2025). YSED: Yemeni Speech Emotion Dataset. Data in Brief. DOI: 10.1016/j.dib.2025.112233. Zenodo: https://zenodo.org/records/15227219.
What's in here
1432 audio clips across… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Yemeni-Speech-Emotion-Dataset.somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.crowd-recital-yi
About
This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.Yadonay-YDX07_Multilingual_Corpus_2026YDX07_Multilingual_Corpus_2026vox2-veri-full
VoxCeleb 2
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
5,994
5,994
118
# of samples
982,808
109,201
36,237
Data Fields
ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
duration (float64): The duration of the segment in seconds.
wav (string): The filepath of the waveform.
start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-full.Yadonay-YDX07_Multilingual_Corpus_2026vox2-veri-3s
VoxCeleb 2
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
5,994
5,994
118
# of samples
982,808
109,201
36,237
Data Fields
ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
duration (float64): The duration of the segment in seconds.
wav (string): The filepath of the waveform.
start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-3s.crowd-whatsapp-yi
About
This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project.
Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot.
Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.vox1-veri-3s
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
299246
33672
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
yoruba-subset
Yoruba Subset Dataset 🎙️🇳🇬
This is a 100-hour subset of the Naija Voices Yoruba speech dataset,
curated by @Abdullah804.
101,391 samples
Converted to 16kHz WAV format
Includes metadata: path, text, speaker_id, gender, age_range, duration
Prepared for ASR/TTS finetuning.
chuvash-asr-final
Chuvash ASR1
Небольшой аудиодатасет для распознавания чувашской речи.
Structure
audio/ — аудиофайлы
train.csv — метаданные с колонками:
file_name
text_chv
client_id
Columns
file_name — относительный путь к аудиофайлу
text_chv — расшифровка на чувашском языке
client_id — идентификатор говорящего
Example
file_name,text_chv,client_id
audio/00001.ogg,Салам,spk01
audio/00002.ogg,Ырӑ ир,spk01
Web demo (single page)
Файл web_app.py… See the full description on the dataset page: https://huggingface.co/datasets/yulia774/chuvash-asr-final.new_datasetmetavoice_librispeech-long_LibriTTS
id
sentence
121_127105_000043_000004
He was handsome and bold and pleasant, offhand and gay and kind.
121_127105_000040_000000
"With this outbreak at last."
121_127105_000012_000001
He passed his hand over his eyes, made a little wincing grimace.
Speaker\Text
121_127105_000043_000004
121_127105_000040_000000
121_127105_000012_000001
7850
🎧
🎧
🎧
6313
🎧
🎧
🎧
422
🎧
🎧
🎧
2086
🎧
🎧
🎧
5895
🎧
🎧
🎧
2803
🎧
🎧
🎧
1919
🎧
🎧
🎧
6319
🎧
🎧
🎧
5536… See the full description on the dataset page: https://huggingface.co/datasets/Tony-Yeh/metavoice_librispeech-long_LibriTTS.CV-Corpus-8.0-2022-01-19-krsl
