datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv-v1.0-segment
CommonVoice v1 Phone-Segment Alignments
Phone-level time alignments for 10 languages of Mozilla Common Voice,
packaged in a canonical segmentation schema with embedded 16 kHz audio. The
phone boundaries come from the charsiu/cv_ali
release of MFA alignments; the audio and transcripts come from
Common Voice Corpus 13.0 (2023-03-09).
Dataset summary
lang
train rows
train hrs
val rows
val hrs
test rows
test hrs
en
1,008,669
1,354.0
3,537
4.9
1,285
1.7
rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.composite_corpus_es_v1.0
Composite dataset for Spanish made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/es: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
openslr: a train split made from the SLR(39,61,67,71,72,73,74,75,108) subsets… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_es_v1.0.composite_corpus_eseu_v1.0
Composite bilingual dataset for Spanish and Basque made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/es: a portion of the "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
mozilla-foundation/common_voice_18_0/eu:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eseu_v1.0.cv-corpus-1.0-en-client_id-grouped
cv-corpus-1.0-en-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.Vaani-Benchmark-V1.0
Vaani-Benchmark-V1.0
A curated ASR evaluation set drawn from the Vaani project. This benchmark contains 5,050 audio segments from 1,103 speakers across 104 Indian districts, each with three independent human transcriptions.
Evaluation Toolkit
A standalone toolkit implementing this benchmark's scoring methodology, plus
Latin-script normalization for code-switched predictions and one-command
publishing of results to a model's HF card, is available at… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0.nigerian-pidgin-1.0
Language:
- Nigerian Pidgin English (West African Pidgin variant)
Dataset Description
Dataset Summary
The Nigerian Pidgin ASR dataset (v1.0) is the first publicly released speech-to-text corpus for Nigerian Pidgin English, a widely spoken lingua franca across Nigeria and West Africa. This dataset comprises over 3,000 audio recordings paired with sentence-level transcriptions, recorded by native speakers across different genders and age groups. It is tailored for… See the full description on the dataset page: https://huggingface.co/datasets/asr-nigerian-pidgin/nigerian-pidgin-1.0.stortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.TWB-Voice-1.0
TWB Voice Dataset v1.0
Dataset Summary
TWB Voice 1.0 is a multilingual speech corpus containing read speech data in three languages from Nigeria: Hausa, Shuwa Arabic, and Kanuri. This dataset was created as part of the TWB Voice project by CLEAR Global (formerly Translators without Borders) to support automatic speech recognition (ASR) development for underrepresented languages.
Languages
Hausa (hau): Major language spoken in Nigeria, Niger, and neighboring… See the full description on the dataset page: https://huggingface.co/datasets/CLEAR-Global/TWB-Voice-1.0.snips_slu_v1.0
Dataset Card for SNIPS SLU v1.0
Dataset Summary
This dataset contains SNIPS SLU Speech Recognition Dataset, available here.
It contains recordings of commands for smart home appliances in English, with info about demographics of the speaker.
ViToSA-1.0
ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances
This is the official repository for the ViToSA 1.0 dataset, introduced in the paper ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances, accepted at Interspeech 2025.The dataset was developed by researchers from the University of Information Technology, VNU-HCM.
Citation Information
If you use this dataset, please cite the following paper:… See the full description on the dataset page: https://huggingface.co/datasets/UIT-ViToSA/ViToSA-1.0.snips_slu_v1.0
Dataset Card for SNIPS SLU v1.0
Dataset Summary
This dataset contains SNIPS SLU Speech Recognition Dataset, available here.
It contains recordings of commands for smart home appliances in English, with info about demographics of the speaker.
