CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MERaLiON /Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching. ASR: Automatic Speech Recognition SQA: Speech Question Answering SDS: Spoken Dialogue Summarization PQA: Paralinguistic Question Answering from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.audio10M<n<100M22 likes28k downloads2y agoHugging Face02cucl2 /AnyAudio-Judge-Corpus AnyAudio-Judge Corpus An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains: An audio clip (referenced relatively under audios/). A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale). A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.audioaudio-text-to-text10K<n<100K0 likes11k downloads2mo agoHugging Face03AudioLLMs /Multitask-National-Speech-Corpus-v1-extendaudio10M<n<100M5 likes6.3k downloads1y agoHugging Face04mort666 /cv_corpus_v22 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. NOTE: currently converting to parquet for convenience.. WIP Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.audioautomatic-speech-recognition1M<n<10M0 likes5.4k downloads9mo agoHugging Face05facebook /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn", `glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M210 likes4.8k downloads10mo agoHugging Face06Serbski-institut /dsb_audio_corpus Acknowledgements Thanks to all speakers that contributed to this dataset! Thanks to "Ludowe Nakładnistwo Domowina" and "Rěčny Centrum WITAJ" for donation of their recordings! audioautomatic-speech-recognition10K<n<100K2 likes2.2k downloads1y agoHugging Face07issai /Kazakh_Speech_Corpus_2 Kazakh Speech Corpus 2 (KSC2) This dataset card describes the KSC2, an industrial-scale, open-source speech corpus for the Kazakh language. Paper: KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus Summary: KSC2 corpus subsumes the previously introduced two corpora: Kazakh Speech Corpus and Kazakh Text-To-Speech 2, and supplements additional data from other sources like tv programs, radio, senate, and podcasts. In total, KSC2 contains around 1.2k hours of high-quality… See the full description on the dataset page: https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2.audioautomatic-speech-recognition10 likes1.9k downloads2y agoHugging Face08benjamin-paine /dinner-party-corpusThis repository contains a reorganized, utterance-focused version of the Dinner Party Corpus, released by Amazon, the Center for Language and Speech Processing (CLSP) and Johns Hopkins University in September 2019. Description The following description is provided in arXiv 1909.13447: We present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment. The corpus was created by recording multiple groups of four Amazon employee… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/dinner-party-corpus.audioautomatic-speech-recognition100K<n<1M1 likes1.7k downloads2y agoHugging Face09TaurenMountain /DAVE-Corpus DAVE-Corpus (Open Subset) Overview DAVE-Corpus (open subset) is a training dataset for blind source separation (BSS) and two-speaker speech separation on Chinese meeting speech. It is the redistributable portion of the training pool of DAVE (arXiv:2608.09288), our system for the ISCSLP 2026 Real-World AVSE Challenge, and is generated end-to-end by the released synthesis pipeline from three permissively licensed corpora — AliMeeting, AISHELL-4 (speech) and MUSAN… See the full description on the dataset page: https://huggingface.co/datasets/TaurenMountain/DAVE-Corpus.audioaudio-to-audio10K<n<100K3 likes1.2k downloads1mo agoHugging Face10ghanaopenai /navigation-corpus-dagbani-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Dag Speech Segments (sentence splitting) 52799 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.audioautomatic-speech-recognition10K<n<100K0 likes978 downloads2mo agoHugging Face11ghanaopenai /navigation-corpus-twi-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Speech Segments (sentence splitting) 52562 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-twi-speech.audioautomatic-speech-recognition10K<n<100K0 likes932 downloads2mo agoHugging Face12rabah2026 /Quran-Ayah-Corpus Quran-Ayah-Corpus: A Multi-Reciter Arabic Quranic Speech Dataset Dataset Description: Ayah-Corpus is a large-scale, multi-reciter Arabic speech dataset meticulously curated for Automatic Speech Recognition (ASR) tasks. It consists of high-quality audio recordings of Quranic verses (Ayahs) paired with their corresponding exact transcriptions. The audio is sourced from two primary repositories: Al-Quran.cloud and EveryAyah.com. This dataset is specifically designed to… See the full description on the dataset page: https://huggingface.co/datasets/rabah2026/Quran-Ayah-Corpus.audioautomatic-speech-recognition100K<n<1M2 likes918 downloads1y agoHugging Face13KathleenKunLiu /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M0 likes727 downloads3mo agoHugging Face14Banglabox /bangla-corpus BanglaBox — Bangladeshi Bangla TTS corpus Anonymous artifact for double-blind review. A Bangladeshi Bangla speech corpus for text-to-speech and zero-shot voice cloning, built with the coverage-driven script pipeline described in the paper (7 domains — news, customer care, teaching, healthcare, e-commerce, finance, IT — with scripts selected under a tiered Jensen–Shannon-divergence objective over phones, diphones, triphones and conjunct clusters (juktakkhor) and filtered by… See the full description on the dataset page: https://huggingface.co/datasets/Banglabox/bangla-corpus.audio100K<n<1M0 likes722 downloads2d agoHugging Face15HiTZ /composite_corpus_es_v1.0 Composite dataset for Spanish made from public available data This dataset is composed of the following public available data: Train split: The train split is composed of the following datasets combined: mozilla-foundation/common_voice_18_0/es: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data) openslr: a train split made from the SLR(39,61,67,71,72,73,74,75,108) subsets… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_es_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes649 downloads1y agoHugging Face16FluidInference /ami-corpus-mirror AMI Corpus Mirror Mirror of the subset of the AMI Meeting Corpus used by FluidAudio diarization benchmarks. Hosted here so CI and local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server (see FluidAudio#752). Contents annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2 (repackaged from the official archive; identical content, including segments/, words/, corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.audio1K<n<10K0 likes623 downloads3mo agoHugging Face17Jzuluaga /atcosim_corpus Dataset Card for ATCOSIM corpus Dataset Summary The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atcosim_corpus.audioautomatic-speech-recognition1K<n<10K19 likes605 downloads4y agoHugging Face18cx-cmu /AgentWebBench-corpus AgentWebBench Corpus Pre-built dense-retrieval corpus for AgentWebBench [ICML 2026], a benchmark for Multi-Agent Coordination in Agentic Web over a realistic 100-website slice of ClueWeb22 (~18.4M documents). This repository holds the embeddings and FAISS indices the benchmark loads at run time, including per-website indices, a global index, and website-level vectors. It does not contain ClueWeb22 text (see Raw documents). Websites: 100 Documents: ~18.4M Embedding dim: 1024… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/AgentWebBench-corpus.audiotext-retrievaln<1K0 likes592 downloads3mo agoHugging Face19murodbek /uzbek-speech-corpus Uzbek Speech Corpus Dataset Summary The Uzbek speech corpus (USC) has been developed in collaboration between ISSAI and the Image and Speech Processing Laboratory in the Department of Computer Systems of the Tashkent University of Information Technologies. The USC comprises 958 different speakers with a total of 105 hours of transcribed audio recordings. To ensure high quality, the USC has been manually checked by native speakers. The USC is primarily designed for… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uzbek-speech-corpus.audioautomatic-speech-recognition100K<n<1M5 likes569 downloads2y agoHugging Face20ghanaopenai /navigation-corpus-ewe-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ewe Speech Segments (sentence splitting) 49348 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-ewe Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-ewe-speech.audioautomatic-speech-recognition10K<n<100K0 likes558 downloads3mo agoHugging Face21abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes539 downloads16d agoHugging Face22atikuwu /karakalpak-speech-corpus 📚 Karakalpak Speech Corpus (107 Hours) The Karakalpak Speech Corpus is the first comprehensive, open-access, community-crowdsourced speech recognition dataset for the Karakalpak language (kaa), a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan (Uzbekistan). Founded and led by Atabek Kadirbergenov alongside a student research team from the Muhammad al-Khwarizmi Specialized School in Nukus, this dataset was created to preserve cultural heritage… See the full description on the dataset page: https://huggingface.co/datasets/atikuwu/karakalpak-speech-corpus.textautomatic-speech-recognition10K<n<100K5 likes505 downloads4d agoHugging Face23yanickschraner /swiss_parliament_corpus Dataset Card for "swiss_parliament_corpus" More Information needed audio10K<n<100K1 likes491 downloads4y agoHugging Face24Akjava /QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-EmotionITA-Corpus Emotion Dataset (100 Japanese Female Voices) 彼のあだ名は言い得て妙だよね 11:A lower-pitched female voice with a strong core ヒューズが飛んだ 100:A slightly quirky female voice that leaves a strong impression Overview This dataset contains 100 female voices generated with Qwen3-TTS. Format: 24kHz mono WAV Source: Link to designed voices About ITA-Corpus Emotion The text is based on the ITA-Corpus Emotion, a public domain dataset containing 100… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-Emotion.audio10K<n<100K4 likes470 downloads8mo agoHugging Face25Jzuluaga /atco2_corpus_1h Dataset Card for ATCO2 test set corpus (1hr set) Dataset Summary ATCO2 project aims at developing a unique platform allowing to collect, organize and pre-process air-traffic control (voice communication) data from air space. This project has received funding from the Clean Sky 2 Joint Undertaking (JU) under grant agreement No 864702. The JU receives support from the European Union’s Horizon 2020 research and innovation programme and the Clean Sky 2 JU members other than… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atco2_corpus_1h.audioautomatic-speech-recognitionn<1K12 likes434 downloads4y agoHugging Face26PleasedPenguin /tpi-va-corpus TPI-VA Corpus TPI-VA Corpus is a speech dataset for studying third-party interruption (TPI) robustness in voice assistants. A TPI setting contains a primary speaker interacting with a voice assistant and a third-party speaker who interrupts before the assistant responds. The dataset is introduced in Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions. The paper frames TPI-awareness as two linked abilities: Discerning speaker… See the full description on the dataset page: https://huggingface.co/datasets/PleasedPenguin/tpi-va-corpus.audio10K<n<100K3 likes423 downloads3mo agoHugging Face27HiTZ /composite_corpus_eseu_v1.0 Composite bilingual dataset for Spanish and Basque made from public available data This dataset is composed of the following public available data: Train split: The train split is composed of the following datasets combined: mozilla-foundation/common_voice_18_0/es: a portion of the "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data) mozilla-foundation/common_voice_18_0/eu:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eseu_v1.0.audioautomatic-speech-recognition100K<n<1M2 likes421 downloads1y agoHugging Face28Tevatron /audiocaps-corpusaudio10K<n<100K0 likes413 downloads1y agoHugging Face29NIVED47 /Sanskrit_ASR_Corpusaudio10K<n<100K1 likes412 downloads2y agoHugging Face30Vack0 /common_voice_corpusThis dataset is processed https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon containing only the english split. Purpose of this repository is faster access for the commonvoice dataset, lower memory required to load the dataset + the option to add it into asr corpus with other datasets. audio1M<n<10M2 likes358 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.