datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus.
MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching.
ASR: Automatic Speech Recognition
SQA: Speech Question Answering
SDS: Spoken Dialogue Summarization
PQA: Paralinguistic Question Answering
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.AnyAudio-Judge-Corpus
AnyAudio-Judge Corpus
An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains:
An audio clip (referenced relatively under audios/).
A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale).
A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.Multitask-National-Speech-Corpus-v1-extendcv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.dsb_audio_corpus
Acknowledgements
Thanks to all speakers that contributed to this dataset!
Thanks to "Ludowe Nakładnistwo Domowina" and "Rěčny Centrum WITAJ" for donation of their recordings!
Kazakh_Speech_Corpus_2
Kazakh Speech Corpus 2 (KSC2)
This dataset card describes the KSC2, an industrial-scale, open-source speech corpus for the Kazakh language.
Paper: KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus
Summary: KSC2 corpus subsumes the previously introduced two corpora: Kazakh Speech Corpus and Kazakh Text-To-Speech 2, and supplements additional data from other sources like tv programs, radio, senate, and podcasts. In total, KSC2 contains around 1.2k hours of high-quality… See the full description on the dataset page: https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2.dinner-party-corpusThis repository contains a reorganized, utterance-focused version of the Dinner Party Corpus, released by Amazon, the Center for Language and Speech Processing (CLSP) and Johns Hopkins University in September 2019.
Description
The following description is provided in arXiv 1909.13447:
We present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment. The corpus was created by recording multiple groups of four Amazon employee… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/dinner-party-corpus.DAVE-Corpus
DAVE-Corpus (Open Subset)
Overview
DAVE-Corpus (open subset) is a training dataset for blind source separation (BSS)
and two-speaker speech separation on Chinese meeting speech. It is the redistributable
portion of the training pool of DAVE
(arXiv:2608.09288), our system for the ISCSLP 2026
Real-World AVSE Challenge, and is generated end-to-end by the released synthesis
pipeline from three permissively licensed corpora — AliMeeting, AISHELL-4
(speech) and MUSAN… See the full description on the dataset page: https://huggingface.co/datasets/TaurenMountain/DAVE-Corpus.navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.navigation-corpus-twi-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Speech Segments (sentence splitting)
52562 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi
Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-twi-speech.Quran-Ayah-Corpus
Quran-Ayah-Corpus: A Multi-Reciter Arabic Quranic Speech Dataset
Dataset Description:
Ayah-Corpus is a large-scale, multi-reciter Arabic speech dataset meticulously curated for Automatic Speech Recognition (ASR) tasks. It consists of high-quality audio recordings of Quranic verses (Ayahs) paired with their corresponding exact transcriptions. The audio is sourced from two primary repositories: Al-Quran.cloud and EveryAyah.com.
This dataset is specifically designed to… See the full description on the dataset page: https://huggingface.co/datasets/rabah2026/Quran-Ayah-Corpus.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.bangla-corpus
BanglaBox — Bangladeshi Bangla TTS corpus
Anonymous artifact for double-blind review. A Bangladeshi Bangla speech corpus for text-to-speech and
zero-shot voice cloning, built with the coverage-driven script pipeline described in the paper
(7 domains — news, customer care, teaching, healthcare, e-commerce, finance, IT — with scripts selected under a
tiered Jensen–Shannon-divergence objective over phones, diphones, triphones and conjunct clusters
(juktakkhor) and filtered by… See the full description on the dataset page: https://huggingface.co/datasets/Banglabox/bangla-corpus.composite_corpus_es_v1.0
Composite dataset for Spanish made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/es: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
openslr: a train split made from the SLR(39,61,67,71,72,73,74,75,108) subsets… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_es_v1.0.ami-corpus-mirror
AMI Corpus Mirror
Mirror of the subset of the AMI Meeting Corpus used by
FluidAudio diarization benchmarks. Hosted here so CI and
local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server
(see FluidAudio#752).
Contents
annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2
(repackaged from the official archive; identical content, including segments/, words/,
corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.atcosim_corpus
Dataset Card for ATCOSIM corpus
Dataset Summary
The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atcosim_corpus.AgentWebBench-corpus
AgentWebBench Corpus
Pre-built dense-retrieval corpus for AgentWebBench [ICML 2026], a benchmark for Multi-Agent Coordination in Agentic Web over a realistic 100-website slice of
ClueWeb22 (~18.4M documents).
This repository holds the embeddings and FAISS indices the benchmark loads at run time, including per-website indices, a global index, and website-level vectors. It does not contain ClueWeb22 text (see Raw documents).
Websites: 100
Documents: ~18.4M
Embedding dim: 1024… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/AgentWebBench-corpus.uzbek-speech-corpus
Uzbek Speech Corpus
Dataset Summary
The Uzbek speech corpus (USC) has been developed in collaboration between ISSAI and the Image and Speech Processing Laboratory in the Department of Computer Systems of the Tashkent University of Information Technologies. The USC comprises 958 different speakers with a total of 105 hours of transcribed audio recordings. To ensure high quality, the USC has been manually checked by native speakers. The USC is primarily designed for… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uzbek-speech-corpus.navigation-corpus-ewe-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ewe Speech Segments (sentence splitting)
49348 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-ewe
Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-ewe-speech.darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.karakalpak-speech-corpus
📚 Karakalpak Speech Corpus (107 Hours)
The Karakalpak Speech Corpus is the first comprehensive, open-access, community-crowdsourced speech recognition dataset for the Karakalpak language (kaa), a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan (Uzbekistan).
Founded and led by Atabek Kadirbergenov alongside a student research team from the Muhammad al-Khwarizmi Specialized School in Nukus, this dataset was created to preserve cultural heritage… See the full description on the dataset page: https://huggingface.co/datasets/atikuwu/karakalpak-speech-corpus.swiss_parliament_corpus
Dataset Card for "swiss_parliament_corpus"
More Information needed
QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-EmotionITA-Corpus Emotion Dataset (100 Japanese Female Voices)
彼のあだ名は言い得て妙だよね
11:A lower-pitched female voice with a strong core
ヒューズが飛んだ
100:A slightly quirky female voice that leaves a strong impression
Overview
This dataset contains 100 female voices generated with Qwen3-TTS.
Format: 24kHz mono WAV
Source: Link to designed voices
About ITA-Corpus Emotion
The text is based on the ITA-Corpus Emotion, a public domain dataset containing 100… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-Emotion.atco2_corpus_1h
Dataset Card for ATCO2 test set corpus (1hr set)
Dataset Summary
ATCO2 project aims at developing a unique platform allowing to collect, organize and pre-process air-traffic control (voice communication) data from air space. This project has received funding from the Clean Sky 2 Joint Undertaking (JU) under grant agreement No 864702. The JU receives support from the European Union’s Horizon 2020 research and innovation programme and the Clean Sky 2 JU members other than… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atco2_corpus_1h.tpi-va-corpus
TPI-VA Corpus
TPI-VA Corpus is a speech dataset for studying third-party interruption (TPI) robustness in voice assistants. A TPI setting contains a primary speaker interacting with a voice assistant and a third-party speaker who interrupts before the assistant responds. The dataset is introduced in Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions.
The paper frames TPI-awareness as two linked abilities:
Discerning speaker… See the full description on the dataset page: https://huggingface.co/datasets/PleasedPenguin/tpi-va-corpus.composite_corpus_eseu_v1.0
Composite bilingual dataset for Spanish and Basque made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/es: a portion of the "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
mozilla-foundation/common_voice_18_0/eu:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eseu_v1.0.audiocaps-corpusSanskrit_ASR_Corpuscommon_voice_corpusThis dataset is processed https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon containing only the english split.
Purpose of this repository is faster access for the commonvoice dataset, lower memory required to load the dataset + the option to add it into asr corpus with other datasets.
