datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.whisper_transcriptions.reazon_speech_all.wer_10.0.vectorizedwhisper_transcriptions.reazon_speech_allpeoples_speech
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.yodas2_sidon
YODAS2-Sidon
Overview
This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks.
We resampled original sidon output to 24kHz due to a storage constraints.
The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus.
MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching.
ASR: Automatic Speech Recognition
SQA: Speech Question Answering
SDS: Spoken Dialogue Summarization
PQA: Paralinguistic Question Answering
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.unsupervised_peoples_speech
Dataset Card for Unsupervised Peoples Speech
Dataset Description
Dataset Summary
The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.
Point of Contact: MLCommons Datasets Discord
Dataset Structure
This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.speech-wikimedia
Dataset Card for Speech Wikimedia
Dataset Summary
The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers.
Each audiofile should have one or more transcriptions in different languages.
Transcription languages
English
German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.gigaspeech
Dataset Card for Gigaspeech
Dataset Description
GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc.
Example Usage
The training split has several configurations of… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech.SpeechRu
Russian Podcasts (unlabeled)
~186k unlabeled Russian-language podcast episodes scraped from the web,
packaged as Parquet shards with the audio bytes embedded. The audio has no
transcripts — this is an unsupervised / self-supervised audio corpus,
suitable for ASR pre-training, speech-representation learning, TTS data
mining, audio classification, and similar tasks.
Each row contains:
audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo),
decoded on-the-fly via the… See the full description on the dataset page: https://huggingface.co/datasets/Sinoosoida/SpeechRu.speech_robust_bench
Dataset Card for "speech_robust_bench"
More Information needed
peoples_speech_v1.0
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech_v1.0.Emotional_SpeechThis dataset contains audio-text pairs in the webdataset format.
The audio files are short speech segments from publicly available videos & the texts are descriptions of emotions the speakers seems to be feeling. Some captions also describe the speakers gender and age.
All files with the substring "part1" in the name contain unique audio files with unique captions.
All files with the substring "part2" , "part3", ... in the name contain the same audio files as in "part1", but with different… See the full description on the dataset page: https://huggingface.co/datasets/EQ4You/Emotional_Speech.rhasspy-speechmls_sidon
MLS-Sidon
Overview
This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
The dataset is provided in WebDataset format for efficient large-scale training.
Source: Multilingual LibriSpeech
Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese
Format: WebDataset (.tar shards)
License: CC-BY-4.0
Dataset Structure
Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.ja_asr.reazon_speech_allLoquaciousSet
LargeScaleASR: 25,000 hours of transcribed and heterogeneous English speech recognition data for research and commercial use.
The full details are available in the paper.
Made of 6 subsets:
large contains 25,000 hours of read / spontaneous and clean / noisy transcribed speech.
medium contains 2,500 hours of read / spontaneous and clean / noisy transcribed speech.
small contains 250 hours of read / spontaneous and clean / noisy transcribed speech.
clean contains 13,000 hours of read… See the full description on the dataset page: https://huggingface.co/datasets/speechbrain/LoquaciousSet.Multitask-National-Speech-Corpus-v1-extendgigaspeech2
Dataset Card for GigaSpeech 2
Dataset Description
GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
Repository: https://github.com/SpeechColab/GigaSpeech2
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.ghana-speech-ipa
Ghana Speech — Audio with IPA Transcripts
Speech with both transcript forms: the original orthography and the IPA phoneme
sequence read off the audio by ASR. Each language is a subset, with real
train/validation splits.
from datasets import load_dataset
ds = load_dataset("ghanaopendata/ghana-speech-ipa", "Akuapem_Twi_twi", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["text"] # original orthography
ds[0]["ipa"] # IPA phonemes
369,347 clips · ~747 h… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech-ipa.ghana-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Language Statistics
Language
Subset
Segments
Duration
Akuapem_Twi
Akuapem_Twi_twi
52,650
63.25h
Anyin
Anyin_any
5,568
13.24h
Asante_Twi
Asante_Twi_twi
143,383
200.02h
Avatime
Avatime_avn
9,956
21.62h
Bassar_Ntcham… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech.legco-speech
香港立法會會議語音數據集
本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。
數據集製作流程
先去香港特別行政區立法會 YouTube下載所有會議紀錄並轉為 16kHz 採樣率嘅 OPUS音頻
用 fsmn-vad 切分所有語音,並用 Qwen3-ASR-1.7B 轉寫成粵文 srt 字幕
轉寫後用正則表達式修正字幕中常見轉寫錯誤
將數據集分成 raw、 segmented 兩個子集傳到HF
子集 subset
raw
segment
總行數 Row number
14,036
9,557,109
總時長 Total duration
22,195.55 hr (79,903,980.00 s)
20471.21 hr (73,696,365.27 s)
平均時長 Average duration
1.58 hr (5692.79 s)
7.71… See the full description on the dataset page: https://huggingface.co/datasets/laubonghaudoi/legco-speech.PromptTSEjapanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.Speech-MASSIVE
Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian, Korean, Dutch, Polish, European Portuguese, Russian, Turkish, and Vietnamese) from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. MASSIVE… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE.multilingual-speech-commands-3lang-raw
Multilingual Speech Commands Dataset (3 Languages, Raw)
This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied.
All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.common_voice_16_0khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech
@inproceedings{moon-etal-2020-beep,
title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection",
author = "Moon, Jihyung and
Cho, Won Ik and
Lee, Junbum",
booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media",
month = jul,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.ghana-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Language Statistics
Language
Subset
Segments
Duration
Akuapem_Twi
Akuapem_Twi_twi
52,650
63.25h
Anyin
Anyin_any
5,568
13.24h
Asante_Twi
Asante_Twi_twi
143,383
200.02h
Avatime
Avatime_avn
9,956
21.62h
Bassar_Ntcham… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech.
