datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Meta_STT_HI_Set1
Meta Speech Recognition Hindi Dataset (Set 1)
This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources.
Dataset Sources and Credits
This dataset combines samples from the following sources:
AI4Bharat Indic Speech Dataset
Source: https://ai4bharat.org/indic-speech-dataset
License: CC-BY 4.0
Citation: Please cite the original paper if you use this data
Common Voice Hindi
Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.Meta_STT_ZH_AIShell3
Meta Speech Recognition Mandarin Dataset (AISHELL3)
This dataset contains both metadata and audio files for Mandarin speech recognition samples from the AISHELL3 corpus.
Dataset Statistics
Splits and Sample Counts
train: 60098 samples
valid: 3163 samples
test: 24772 samples
Example Samples
train
{
"audio_filepath": "/external4/datasets/Mandarin/AISHELL3/wavs_train/SSB00430356.wav",
"text": "她以 ENTITY_PRODUCT 滴鸡精 END 调养身体。 AGE_14_25… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_ZH_AIShell3.Meta_STT_EN_Set2
Meta Speech Recognition English Dataset (Set 2)
This dataset contains both metadata and audio files for English speech recognition samples.
Dataset Statistics
Splits and Sample Counts
train: 42961 samples
valid: 2387 samples
test: 2387 samples
Example Samples
train
{
"audio_filepath": "/external1/datasets/asr-himanshu/avspeech-data/audio/AzSutepklXI_2.wav",
"text": "To Jesus, so God is faithful, because when he keeps, you know, when… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set2.betrac-2026-with-meta
betrac-2026-with-meta
Annotated speech dataset created with Whissle Annotator — a multimodal annotation pipeline for speech, NLP, and visual analysis.
Source Dataset
This dataset is derived from the following HuggingFace dataset(s):
BeTraC/betrac-2026:0:50
BeTraC/betrac-2026:50:50
BeTraC/betrac-2026:100:50
BeTraC/betrac-2026:150:50
BeTraC/betrac-2026:200:50
BeTraC/betrac-2026:250:50
BeTraC/betrac-2026:300:50
BeTraC/betrac-2026:350:50
BeTraC/betrac-2026:400:50… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/betrac-2026-with-meta.Meta_STT_EURO_Set1
Meta Speech Recognition European Languages Dataset (v1)
This dataset contains only the metadata (JSON/Parquet) for European language speech recognition samples.Audio files are NOT included.
Data Download Links
CommonVoice
CommonVoice Dataset
German (de)
English (en)
Spanish (es)
French (fr)
Italian (it)
Portuguese (pt)
Multilingual LibriSpeech (MLS)
Multilingual LibriSpeech Dataset
German: mls_german.tar.gz
English: mls_english.tar.gz
Spanish:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EURO_Set1.Meta_STT_EN_Set1
Meta Speech Recognition English Dataset (v1)
This dataset contains only the metadata (JSON/Parquet) for English speech recognition samples.Audio files are NOT included.
Data Download Links
CommonVoice: https://commonvoice.mozilla.org/en/datasets
People's Speech: https://huggingface.co/datasets/MLCommons/peoples_speech
LibriSpeech:
train-clean-100
train-clean-360
train-other-500
Dataset Statistics
Splits and Sample Counts
train: 2338349… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set1.Meta_STT_SLAVIC_CommonVoice
Meta Speech Recognition Slavic Languages Dataset (Common Voice)
This dataset contains metadata for Slavic language speech recognition samples from Common Voice.
Dataset Sources and Credits
This dataset contains samples from Mozilla Common Voice:
Source: https://commonvoice.mozilla.org/en/datasets
License: CC0-1.0
Citation: Please acknowledge Mozilla Common Voice if you use this data
Languages Included
The dataset includes the following Slavic languages:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_SLAVIC_CommonVoice.emilia-ja-plus-metadata
Emilia Dataset JA Plus — normalized metadata
This is an audit-backed, metadata-only derivative of
ayousanz/Emilia-Dataset-JA-Plus. It does not bundle audio payloads.
Verified snapshot statistics
Metric
Value
Metadata rows
78,748
Unique IDs
78,748
Duplicate IDs
0
Unique speakers
10,046
Duration represented by metadata
145.06 hours
Language labels
ja: 78,748
Unique transcripts
64,034
Duplicate transcript rows
14,714
Japanese-labelled rows… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/emilia-ja-plus-metadata.voicevox-voice-corpus-metadata
VOICEVOX synthetic Japanese speech corpus
This metadata derivative documents ayousanz/voicevox-voice-corpus at
39dff6b254bf3118accad04446238ae3147156a6.
Audio payloads were not downloaded during this metadata audit.
Verified repository inventory
Corpus
Voice/style directories
WAV files
WAV bytes
ROHAN-corpus
87
400,200
90.16 GB
ita-corpus
87
36,888
6.64 GB
tsukuyomi-chan-corpus
87
8,700
3.07 GB
Total WAV paths: 445,788
Total WAV bytes: 99.87… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/voicevox-voice-corpus-metadata.Meta_STT_MADASR2.0_train_lg
Dataset Card for Dataset Name
https://sites.google.com/view/respinasrchallenge2025/home?authuser=0
This is the trainig dataset provided for track-3 and track-4 of this task. We enhance the dataset with entity tagging, emotion, age, gender and intent.
WhissleAI participated in this challenge.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_MADASR2.0_train_lg.Meta_STT_EN-IN_Tech_Interviews
Meta STT EN-IN Tech Interviews
Indian English speech recognition dataset sourced from technical interviews, annotated with rich speech metadata including age group, gender, emotion, and intent. Designed for training multi-task ASR models that jointly predict transcriptions and speaker attributes.
Dataset Details
Property
Value
Train examples
58,000
Validation examples
1,204
Language
English (Indian accent)
Audio
16 kHz
Total size
~28 GB… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN-IN_Tech_Interviews.
