CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face02fsicoli /common_voice_22_0 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_22_0.automatic-speech-recognition100B<n<1T20 likes28k downloads1y agoHugging Face03Reverb /voxceleb2 VoxCeleb2 Dataset This is the VoxCeleb2 dataset, a large-scale speaker identification dataset. Dataset Description VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube. Files vox2_dev_mp4_part*: Multipart archive containing MP4 video files vox2_dev_txt: Text files with speaker/utterance metadata vox2_meta.csv: Dataset metadata Usage To extract the multipart archive: # Using 7zip 7z x… See the full description on the dataset page: https://huggingface.co/datasets/Reverb/voxceleb2.automatic-speech-recognition100K<n<1M22 likes7.9k downloads1y agoHugging Face04KBLab /rixvox-v2 RixVox-v2: A Swedish parliamentary speech dataset RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.audioautomatic-speech-recognition1M<n<10M12 likes7k downloads1y agoHugging Face05anuj-inavlabs /Thinkspark-v2-270m-training-data ThinkSpark-v2-350M — training data Full-duplex floor-controller (Section 8) training corpus: playable audio + text, paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain, gender, prosody, agent text) and Soniox character-level timestamps. Dataset Viewer Default split is parquet with a real Audio feature — a player renders inline next to the text in the Hub UI: column type description audio Audio playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.text-to-speech1K<n<10K0 likes6.9k downloads21d agoHugging Face06ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K6 likes6.4k downloads7mo agoHugging Face07takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.1k downloads6mo agoHugging Face08speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes6k downloads6mo agoHugging Face09mort666 /cv_corpus_v22 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. NOTE: currently converting to parquet for convenience.. WIP Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.audioautomatic-speech-recognition1M<n<10M0 likes5.4k downloads9mo agoHugging Face10joujiboi /japanese-anime-speech-v2 Japanese Anime Speech Dataset V2 日本語はこちら japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models. The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels. This dataset is not an updated version of japanese-anime-speech-v1. For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset. The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.audioautomatic-speech-recognition100K<n<1M152 likes4.1k downloads11mo agoHugging Face11Centi234 /WorldSpeech WorldSpeech A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Centi234/WorldSpeech.audioautomatic-speech-recognition10M<n<100M2 likes3.9k downloads4mo agoHugging Face12oddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M29 likes3.7k downloads2mo agoHugging Face13zhifeixie /StreamAudio-2M StreamAudio-2M Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips are organised into six task subsets. Subsets Subset Rows Description Stream_Audio_Understanding 90,738 Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA Real_time_ASR 28,109 Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.tabularaudio-classification100K<n<1M30 likes2.8k downloads4mo agoHugging Face14zhifeixie /Voices-in-the-Wild-2M Voices in the Wild Project Page | Paper | GitHub Voices in the Wild (Voices-in-the-Wild-2M) is a large-scale automatic speech recognition (ASR) dataset designed for robustness training and evaluation under diverse, real-world acoustic conditions. It covers 7 classic acoustic phenomena (including noise, far-field speech, obstruction, echo/reverberation, recording artifacts, electronic distortion, and transmission dropout) and 54 physically plausible compound scenarios. The… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/Voices-in-the-Wild-2M.audioautomatic-speech-recognition50 likes2.8k downloads4mo agoHugging Face15intronhealth /afrispeech-200AFRISPEECH-200 is a 200hr Pan-African speech corpus for clinical and general domain English accented ASR; a dataset with 120 African accents from 13 countries and 2,463 unique African speakers. Our goal is to raise awareness for and advance Pan-African English ASR research, especially for the clinical domain.automatic-speech-recognition10K<n<100K40 likes2.4k downloads3y agoHugging Face16MushanW /GLOBE_V2 Important notice Differences between V2 version and the version described in paper: The V2 version provide audio in 44.1kHz sample rate. (Supersampling) The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues. Globe The full paper can be accessed here: arXiv An online demo can be accessed here: Github Abstract This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the… See the full description on the dataset page: https://huggingface.co/datasets/MushanW/GLOBE_V2.audiotext-to-audio100K<n<1M15 likes2.4k downloads2y agoHugging Face17ghanaopenai /ghana-english-asr-2700hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. 🇬🇭 Ghana English ASR Dataset A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Automatic Speech Recognition (ASR) models on West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.audioautomatic-speech-recognition100K<n<1M7 likes2.2k downloads3mo agoHugging Face18overflowwwww /yt-danish-public-v2audioaudio-classification100K<n<1M0 likes2.2k downloads2y agoHugging Face19facebook /2M-Belebele 2M-Belebele Highly-Multilingual Speech and American Sign Language Comprehension Dataset We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL). The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.tabularquestion-answering10K<n<100K13 likes2.2k downloads2y agoHugging Face20MohamedRashad /SADA22 Dataset Card for SADA (Saudi Audio Dataset for Arabic) Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.audioautomatic-speech-recognition100K<n<1M30 likes2.1k downloads1y agoHugging Face21Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes2k downloads3mo agoHugging Face22kensho /SPGISpeech2.0gated Dataset Card for SPGISpeech 2.0 Dataset Details Dataset Overview We are excited to present SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR).… See the full description on the dataset page: https://huggingface.co/datasets/kensho/SPGISpeech2.0.audioautomatic-speech-recognition100K<n<1M4 likes1.9k downloads5mo agoHugging Face23issai /Kazakh_Speech_Corpus_2 Kazakh Speech Corpus 2 (KSC2) This dataset card describes the KSC2, an industrial-scale, open-source speech corpus for the Kazakh language. Paper: KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus Summary: KSC2 corpus subsumes the previously introduced two corpora: Kazakh Speech Corpus and Kazakh Text-To-Speech 2, and supplements additional data from other sources like tv programs, radio, senate, and podcasts. In total, KSC2 contains around 1.2k hours of high-quality… See the full description on the dataset page: https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2.audioautomatic-speech-recognition10 likes1.9k downloads2y agoHugging Face24Theafricatechguy /common_voice_21_0 Dataset Card for Common Voice Corpus 21.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash… See the full description on the dataset page: https://huggingface.co/datasets/Theafricatechguy/common_voice_21_0.automatic-speech-recognition100B<n<1T0 likes1.9k downloads26d agoHugging Face25ivrit-ai /audio-v2gatedThis dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It wa released on April 20th, 2025. You can find the full list of sources in this dataset under the dataset's sources.txt. Paper: https://arxiv.org/abs/2307.08720 If you use our datasets, the following quote is preferable: @misc{marmor2023ivritai, title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development}, author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2.audio-classification10K<n<100K3 likes1.8k downloads8mo agoHugging Face26disco-eth /EuroSpeech-24kHz EuroSpeech 24 kHz Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. Dataset Summary Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.audioautomatic-speech-recognition10M<n<100M3 likes1.7k downloads5mo agoHugging Face27facebook /2M-Flores-ASL 2M-Flores As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest sentences in the original flores200 dataset. To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded. The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time. The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.tabulartranslation1K<n<10K2 likes1.7k downloads2y agoHugging Face28fsicoli /common_voice_21_0 Dataset Card for Common Voice Corpus 21.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_21_0.automatic-speech-recognition100B<n<1T2 likes1.4k downloads1y agoHugging Face29shb777 /gemini-flash-2.0-speech 🎙️ Gemini Flash 2.0 Speech Dataset This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English. 🏅 #1 Trending Audio Dataset in Feb 2025 🏅 Used in training of Kokoro TTS and LLaSA 1B 〽️ Stats Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours) Average duration: 10.83 seconds Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.audiotext-to-speech10K<n<100K60 likes1.4k downloads1y agoHugging Face30suleiman2003 /afri-temp-data4 AfricanVoices Hausa -- Train Split Hausa speech dataset from AfricanVoices.io. Usage from datasets import load_dataset ds = load_dataset("suleiman2003/afri-temp-data4", split="train") print(ds[0]) # {'audio': Audio(...), 'transcript': '...', 'gender': '...', ...} Structure Each batch is in its own subdirectory under train/: train/ batch_1/ *.flac + metadata.csv batch_2/ *.flac + metadata.csv ... Audio files are FLAC format. Metadata… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/afri-temp-data4.audioautomatic-speech-recognition1K<n<10K0 likes1.1k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.