CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ARTPARK-IISc /VaanigatedVAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.audioautomatic-speech-recognition1M<n<10M157 likes19k downloads9d agoHugging Face02OmniAICreator /ASMR-Archive-Processed ASMR-Archive-Processed (WIP) Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended. Work in Progress — expect breaking changes while the pipeline and data layout stabilize. This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.imageautomatic-speech-recognition98 likes8.5k downloads6mo agoHugging Face03ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K6 likes6.9k downloads7mo agoHugging Face04oddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M29 likes4.5k downloads2mo agoHugging Face05aranemini /central-kurdish-pseudolabel Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.audioautomatic-speech-recognition1M<n<10M2 likes4.2k downloads3mo agoHugging Face06kadirnar /voicehub-arena-seed-tts-eval VoiceHub Arena — full English Seed-TTS-Eval 35,904 synthesized WAV files: 33 model families × the same 1,088 target texts. The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB. All 198 shards and every WAV SHA256 were verified after backup. Interactive leaderboard and all audio samples · Source repository (access required). Contents audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.audiotext-to-speech10K<n<100K0 likes2.8k downloads10d agoHugging Face07MohamedRashad /MASC-Arabic MASC Arabic Dataset Card Dataset Summary MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels. The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MASC-Arabic.audioautomatic-speech-recognition100K<n<1M8 likes2.8k downloads6mo agoHugging Face08moaead /dialectal-arabic-voices Dialectal Arabic Voices An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps). 59,812 recordings · approximately 10,634.4 hours · 556.22 GB Column Description audio Original audio, embedded in the Parquet file transcript_text Empty; ASR transcripts are stored in a separate private dataset language Dialect code: ps (Palestinian) source Original channel or account name Audio retains its… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.audioautomatic-speech-recognition10K<n<100K0 likes1.9k downloads1h agoHugging Face09oddadmix /arabic-audio-collection-algerian-loubna-stories Loubna Stories Arabic Speech Dataset Dataset Summary The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.audiotext-to-speech10K<n<100K0 likes1.9k downloads3mo agoHugging Face10ARTPARK-IISc /Vaani-transcription-partgatedThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages. This table represents the audio and transcription duration data for various languages. Language Angami Angika Ao Assamese Awadhi Bajjika Bearybashe Bengali Bhili Bhojpuri Bundeli Chakhesang Chakma Chhattisgarhi English Garhwali Garo Gondi Gujarati Halbi Haryanvi Hindi IduMishmi Kannada Kashmiri Karbi Khariboli Khortha Kokborok Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.audioautomatic-speech-recognition1M<n<10M20 likes1.5k downloads6mo agoHugging Face11linagora /linto-dataset-audio-ar-tn LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task This is the first packaged version of the datasets used to train the Linto Tunisian dialect with code-switching STT (linagora/linto-asr-ar-tn). Dataset Summary Dataset composition Sources Data Table Data sources Content Types Languages and Dialects Example use (python) License Citations Dataset Summary The LinTO DataSet Audio for Arabic Tunisian is a diverse… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn.audioautomatic-speech-recognition10K<n<100K22 likes1.5k downloads1y agoHugging Face12MohamedRashad /mgb2-arabic MGB-2: Arabic Multi-Dialect Broadcast Media Recognition Dataset Description Dataset Summary The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/mgb2-arabic.audioautomatic-speech-recognition100K<n<1M7 likes1.3k downloads9mo agoHugging Face13ArtificialAnalysis /VoxPopuli-Cleaned-AA VoxPopuli-Cleaned-AA Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models. This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes1.1k downloads7mo agoHugging Face14Dr-AliGomaa /ar-quran-hadith14books-MSA ar-quran-hadith14books-MSA Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general Modern Standard Arabic, under one construction pipeline and one text convention. ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter the meaning of scripture, and because chatbots, search and summarizers increasingly answer from transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.audioautomatic-speech-recognition10K<n<100K6 likes1.1k downloads1mo agoHugging Face15Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes888 downloads4mo agoHugging Face16aranemini /northern-kurdish-raw-audio Northern Kurdish Raw Audio Collection Overview This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources. The collection was assembled to support research and development in: Automatic Speech Recognition (ASR) Speech Translation (ST) Text-to-Speech (TTS) Self-supervised Learning (SSL) Spoken Language Understanding (SLU) The dataset contains more than 2,000 hours… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-raw-audio.audioautomatic-speech-recognition1K<n<10K0 likes754 downloads3mo agoHugging Face17MBZUAI /ArVoice ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Hawau Olamide Toyin, Rufael Marew, Humaid Alblooshi, Samar M. Magdy, Hanan Aldarmaki {hawau.toyin, hanan.aldarmaki}@mbzuai.ac.ae ArVoice is a multi-speaker Modern Standard Arabic (MSA) speech corpus with fully diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArVoice.audiotext-to-speech10K<n<100K33 likes746 downloads11mo agoHugging Face18noxwano /ASMR-Archive-Processed-SFW ASMR-Archive-Processed-SFW Overview This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset. We filtered the original dataset to include only records where the nsfw metadata flag is false. To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled. The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.audioautomatic-speech-recognition1M<n<10M9 likes744 downloads6mo agoHugging Face19ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes697 downloads3mo agoHugging Face20linagora /linto-dataset-audio-ar-tn-augmented LinTO DataSet Audio for Arabic Tunisian Augmented A collection of Tunisian dialect audio and its annotations for STT task This is the augmented datasets used to train the Linto Tunisian dialect with code-switching STT linagora/linto-asr-ar-tn. Dataset Summary Dataset composition Sources Content Types Languages and Dialects Example use (python) License Citations Dataset Summary The LinTO DataSet Audio for Arabic Tunisian Augmented is a dataset that builds on LinTO… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn-augmented.audioautomatic-speech-recognition100K<n<1M7 likes675 downloads1y agoHugging Face21davidggphy /librispeech-arpabet-processed LibriSpeech ARPAbet Processed Dataset Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss. Dataset Description This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing: Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary Filtering: Samples with out-of-vocabulary words (not in CMU… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.audioautomatic-speech-recognition10K<n<100K0 likes528 downloads8mo agoHugging Face22MohamedRashad /common-voice-18-arabic Dataset Card for Common Voice 18 – Arabic Edition Dataset Summary This dataset is an unofficial Arabic-only extraction of Mozilla Common Voice Corpus 18.0, prepared for Automatic Speech Recognition (ASR) research and development. It is derived from the original Common Voice 18 release and filtered to include Arabic (ar) speech data only, while preserving the original dataset structure, splits, and metadata fields. The dataset consists of validated, unvalidated, and… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/common-voice-18-arabic.audioautomatic-speech-recognition100K<n<1M5 likes475 downloads9mo agoHugging Face23deepghs /arknights_voices_zh ZH Voice-Text Dataset for Arknights Waifus This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models. Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset. 12431 records, 25.9 hours in total. Average duration is approximately 7.49s. id char_id voice_actor_name voice_title voice_text time sample_rate file_size filename mimetype file_url char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.tabularautomatic-speech-recognition10K<n<100K6 likes432 downloads2y agoHugging Face24arda-argmax /simchoir-parquet FastMSS synthetic multi-speaker meetings - parquet edition Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring. Subsets and splits debug — splits: train — 1 mixtures, 1.6 min total, 6 unique speakers… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/simchoir-parquet.audioautomatic-speech-recognition100K<n<1M0 likes362 downloads5mo agoHugging Face25xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes359 downloads2y agoHugging Face26oddadmix /arabic-audio-collection-algerian-kahwa-postcast Kahwa Postcast Arabic Speech Dataset Dataset Summary The Kahwa Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 110 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-kahwa-postcast.audiotext-to-speech10K<n<100K2 likes356 downloads3mo agoHugging Face27Chillarmo /common_voice_20_armenian Common Voice 20 - Armenian This dataset is the Armenian portion of Mozilla's Common Voice 20.0 release, a massively multilingual collection of transcribed speech intended for speech technology research and development. Dataset Details Language: Armenian (hy) Source: Mozilla Common Voice Version: 20.0 License: CC0-1.0 audioautomatic-speech-recognition10K<n<100K1 likes351 downloads1y agoHugging Face28KeisukeMiyamoto /nhk-archive-audio-30sgated NHK Archives Audio 30s This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection. The dataset contains 344,722 accepted clips, totaling 1,661.01 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description. This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.audioautomatic-speech-recognition100K<n<1M0 likes342 downloads8h agoHugging Face29Archime /french_tv_media_dataset_2026 Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus Résumé (Abstract) Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.audioautomatic-speech-recognition10K<n<100K4 likes328 downloads8mo agoHugging Face30ymoslem /CoVoST2-EN-AR Dataset Description CoVoST 2 is a large-scale multilingual speech translation corpus based on Common Voice, developed by FAIR. This is the English-to-Arabic portion of the dataset. The original dataset can be found here. Data Splits (EN-AR) lang train validation test EN-AR 289430 15531 15531 AR-EN 2283 1758 1695 Citation @misc{wang2020covost, title={CoVoST 2: A Massively Multilingual Speech-to-Text Translation Corpus}, author={Changhan… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/CoVoST2-EN-AR.audioautomatic-speech-recognition100K<n<1M5 likes317 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.