CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes48k downloads2y agoHugging Face02zaibihassan /Quranic-Recitation-Data 🌟 Overview Quranic Recitation Dataset (Word-by-Word Sync) is a highly optimized, production-ready dataset containing high-quality audio recitations of the Holy Quran synchronized at the word-by-word level. This dataset features 135 world-renowned reciters, with every Surah (114 chapters) mapped precisely to millisecond-accurate word timestamps. It is designed for modern Islamic mobile and web applications — served via a Cloudflare Edge CDN with native… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data.audioautomatic-speech-recognition10K<n<100K5 likes22k downloads3h agoHugging Face03legacy-datasets /common_voiceCommon Voice is Mozilla's initiative to help teach machines how real people speak. The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.automatic-speech-recognition100K<n<1M148 likes15k downloads2y agoHugging Face04ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads19h agoHugging Face05sanganaka /Vedavani-Dataset Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry Vedavani is the first benchmark dataset for automatic speech recognition (ASR) on Vedic Sanskrit poetry, consisting of richly annotated verses from the Rig Veda and Atharva Veda. This corpus captures the unique prosodic structure, phonetic complexity, and chanting style found in traditional Vedic recitation. 🔗 Paper: Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry (ACL 2025)📁 GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/Vedavani-Dataset.audioautomatic-speech-recognition0 likes10k downloads1y agoHugging Face06Digital-Divide-Data /Luhya-ASR-Data-subset-642H Luhya ASR Data Subset 642H Luhya speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M1 likes8.4k downloads1mo agoHugging Face07LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B89 likes7.1k downloads6mo agoHugging Face08anuj-inavlabs /Thinkspark-v2-270m-training-data ThinkSpark-v2-350M — training data Full-duplex floor-controller (Section 8) training corpus: playable audio + text, paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain, gender, prosody, agent text) and Soniox character-level timestamps. Dataset Viewer Default split is parquet with a real Audio feature — a player renders inline next to the text in the Hub UI: column type description audio Audio playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.text-to-speech1K<n<10K0 likes6.9k downloads21d agoHugging Face09espnet /Bagpiper_SFT_Data Bagpiper SFT Data Release status: the validated Parquet release is being uploaded. The homepage and metadata may appear before every large shard is committed. Bagpiper SFT Data is the supervised fine-tuning corpus for Bagpiper, an open-ended audio language model that understands and generates speech, music, environmental sound, and their mixtures through rich textual captions and planning. The public release has exactly two configurations: Configuration Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.audioaudio-classification1M<n<10M1 likes6.5k downloads2mo agoHugging Face10ehabnegm /100-hour-Egyptian-dataset-single-speaker Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts. Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data. Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home. This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.audiotext-to-speech10K<n<100K11 likes6.2k downloads2mo agoHugging Face11XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.8k downloads5mo agoHugging Face12zaibihassan /Quranic-Translation-Audio-Data Overview Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories. Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.audioaudio-to-audio1K<n<10K1 likes4.5k downloads3mo agoHugging Face13zaibihassan /Quranic-Word-By-Word-Audio-Data 🌟 Overview Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines: Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening. Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow. Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.audioautomatic-speech-recognition100K<n<1M4 likes4.1k downloads4mo agoHugging Face14Digital-Divide-Data /Somali-ASR-Subset-68H Somali ASR Subset 68H Somali speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M3 likes4k downloads1mo agoHugging Face15Digital-Divide-Data /khmer-speech-dataset Khmer ASR Cultural Dataset 727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.audioautomatic-speech-recognition100K<n<1M26 likes3.4k downloads3mo agoHugging Face16ayousanz /Emilia-Dataset-JA-Plus Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2024/08/28: Welcome to join Amphion's Discord channel to stay connected and engage with our community! 2024/08/27: The Emilia dataset is now publicly available! Discover the most extensive and diverse speech generation dataset with… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/Emilia-Dataset-JA-Plus.text-to-speech10M<n<100M1 likes3.4k downloads2y agoHugging Face17nithinraok /asr-leaderboard-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.audioautomatic-speech-recognition100K<n<1M4 likes3k downloads1y agoHugging Face18Digital-Divide-Data /Kamba-ASR-Data-Subset-484H Kamba ASR Data Subset 484H Kamba speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M0 likes2.9k downloads1mo agoHugging Face19espnet /Bagpiper_PreTrain_Data Bagpiper Pretraining Data Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated with Bagpiper, an open-ended audio language model that learns bidirectional mappings between audio and comprehensive text descriptions across speech, music, environmental sound, and mixtures. The en metadata describes the primary rich-caption language. Source audio can contain speech or singing in other languages; it is not an English-only audio guarantee. The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.tabularautomatic-speech-recognition10K<n<100K0 likes2.9k downloads2mo agoHugging Face20opedromartins /ASR-datasets-ptbr 📚 Datasets de Áudio em Português (PT-BR) Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition). O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade. 📂 Datasets Integrados A tabela abaixo lista todos os datasets incluídos, com suas informações: Dataset Config Name TOTAL train test validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.audioautomatic-speech-recognition1M<n<10M13 likes2.5k downloads1y agoHugging Face21bhyuan /gptsovits_datasetgated bhyuan/gptsovits_dataset GPT-SoVITS speech dataset, packed as WebDataset tar shards. Layout data/ train/ metadata.csv audio/ train-000.tar train-001.tar ... validation/ metadata.csv audio/ validation-000.tar ... test/ metadata.csv audio/ test-000.tar ... Shard counts: youshengshu_v5_test: 6536 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.audiotext-to-speech10M<n<100M1 likes2.5k downloads4mo agoHugging Face22Digital-Divide-Data /Gusii-ASR-Data-Subset-470H Gusii ASR Data Subset 470H Gusii speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M0 likes2k downloads1mo agoHugging Face23UBC-NLP /SimbaBench_dataset SibmaBench Data Release & Benchmarking To evaluate your model on SimbaBench across all supported tasks (ASR, TTS, and SLID), simply load the corresponding configuration for the task and language you wish to benchmark. Each task is organized by configuration name (e.g., asr_test_afr, tts_test_wol, slid_61_test). Loading a configuration provides the standardized evaluation split for that specific benchmark.Example: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/SimbaBench_dataset.audioautomatic-speech-recognition100K<n<1M0 likes1.9k downloads7mo agoHugging Face24ken-sungmin /propagator-multimodal-pretraining-data Propagator Multimodal Pretraining Data This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format. This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout. Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.texttext-generation0 likes1.7k downloads3mo agoHugging Face25Aratako /LiquidAI-Hackathon-Tokyo-CPT-Data LiquidAI-Hackathon-Tokyo-CPT-Data Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。 automatic-speech-recognition1M<n<10M6 likes1.5k downloads1y agoHugging Face26linagora /linto-dataset-audio-ar-tn LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task This is the first packaged version of the datasets used to train the Linto Tunisian dialect with code-switching STT (linagora/linto-asr-ar-tn). Dataset Summary Dataset composition Sources Data Table Data sources Content Types Languages and Dialects Example use (python) License Citations Dataset Summary The LinTO DataSet Audio for Arabic Tunisian is a diverse… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn.audioautomatic-speech-recognition10K<n<100K22 likes1.3k downloads1y agoHugging Face27DigiGreen /Agri_STT_Benchmarking_Dataset Agri STT Benchmarking Dataset 10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository. Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.audioautomatic-speech-recognition10K<n<100K1 likes1.2k downloads2mo agoHugging Face28ekacare /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K16 likes1.2k downloads1y agoHugging Face29HKUSTAudio /Audio-FLAN-Datasetgated Audio-FLAN Dataset (Paper) (the FULL audio files and jsonl files are still updating) An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound. 1. Dataset Structure The Audio-FLAN-Dataset has the following directory structure: Audio-FLAN-Dataset/ ├── audio_files/ │ ├── audio/ │ │ └── 177_TAU_Urban_Acoustic_Scenes_2022/ │ │ └── 179_Audioset_for_Audio_Inpainting/ │ │ └── ... │ ├── music/ │ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.audiotext-to-speech10M<n<100M47 likes1.2k downloads1y agoHugging Face30Digital-Divide-Data /khm-asr-cultural Khmer ASR Cultural Dataset 134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.audioautomatic-speech-recognition10K<n<100K9 likes1.1k downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.