datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Meta_STT_HI_Set1
Meta Speech Recognition Hindi Dataset (Set 1)
This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources.
Dataset Sources and Credits
This dataset combines samples from the following sources:
AI4Bharat Indic Speech Dataset
Source: https://ai4bharat.org/indic-speech-dataset
License: CC-BY 4.0
Citation: Please cite the original paper if you use this data
Common Voice Hindi
Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.Meta_STT_EN_Set2
Meta Speech Recognition English Dataset (Set 2)
This dataset contains both metadata and audio files for English speech recognition samples.
Dataset Statistics
Splits and Sample Counts
train: 42961 samples
valid: 2387 samples
test: 2387 samples
Example Samples
train
{
"audio_filepath": "/external1/datasets/asr-himanshu/avspeech-data/audio/AzSutepklXI_2.wav",
"text": "To Jesus, so God is faithful, because when he keeps, you know, when… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set2.primewords_chinese_corpus_set_1YouTube-Evaluation-Set
Awaaz se Alfaaz — YouTube Evaluation Set
This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.openslr-32-hq-SA-languages-Setswana
High quality TTS data for four South African languages - Setswana
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Setswana
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.Audio-Understanding-Test-Set
Audio Understanding Test Set
A structured dataset for evaluating audio understanding capabilities of multimodal AI models. Contains 137 test prompts across 22 categories, paired with a 20-minute voice sample and 49 completed model outputs from Gemini 3.1 Flash Lite.
Overview
Property
Value
Total prompts
137
Implemented (with prompt text)
49
Suggested (description only)
88
Completed outputs
49
Categories
22
Model under test… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Test-Set.asr_evaluate_set
Speech-to-Text Evaluation Dataset
Dataset Overview
This dataset is designed for evaluating Uzbek speech-to-text (STT) models on real-world conversational speech data. The audio samples were collected from various open Telegram groups, capturing natural voice messages in diverse acoustic conditions and speaking styles.
Key Statistics
Total Samples: 745 audio files
Total Duration: 1 hour 40 minutes (~100 minutes)
Average Duration: ~8 seconds per sample
Source:… See the full description on the dataset page: https://huggingface.co/datasets/OvozifyLabs/asr_evaluate_set.golden-eval-set
Golden Eval Set
Private Vietnamese audio evaluation set with audio and transcription columns.
asr_evaluate_set
Speech-to-Text Evaluation Dataset
Dataset Overview
This dataset is designed for evaluating Uzbek speech-to-text (STT) models on real-world conversational speech data. The audio samples were collected from various open Telegram groups, capturing natural voice messages in diverse acoustic conditions and speaking styles.
Key Statistics
Total Samples: 745 audio files
Total Duration: 1 hour 40 minutes (~100 minutes)
Average Duration: ~8 seconds per sample
Source:… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/asr_evaluate_set.Meta_STT_EURO_Set1
Meta Speech Recognition European Languages Dataset (v1)
This dataset contains only the metadata (JSON/Parquet) for European language speech recognition samples.Audio files are NOT included.
Data Download Links
CommonVoice
CommonVoice Dataset
German (de)
English (en)
Spanish (es)
French (fr)
Italian (it)
Portuguese (pt)
Multilingual LibriSpeech (MLS)
Multilingual LibriSpeech Dataset
German: mls_german.tar.gz
English: mls_english.tar.gz
Spanish:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EURO_Set1.Meta_STT_EN_Set1
Meta Speech Recognition English Dataset (v1)
This dataset contains only the metadata (JSON/Parquet) for English speech recognition samples.Audio files are NOT included.
Data Download Links
CommonVoice: https://commonvoice.mozilla.org/en/datasets
People's Speech: https://huggingface.co/datasets/MLCommons/peoples_speech
LibriSpeech:
train-clean-100
train-clean-360
train-other-500
Dataset Statistics
Splits and Sample Counts
train: 2338349… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set1.nchlt_speech_setswana
NCHLT Speech Corpus -- Setswana
This is the Setswana language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): tsn
URI: https://hdl.handle.net/20.500.12185/281
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/nchlt_speech_setswana.
