CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zaibihassan /Quranic-Recitation-Data 🌟 Overview Quranic Recitation Dataset (Word-by-Word Sync) is a highly optimized, production-ready dataset containing high-quality audio recitations of the Holy Quran synchronized at the word-by-word level. This dataset features 135 world-renowned reciters, with every Surah (114 chapters) mapped precisely to millisecond-accurate word timestamps. It is designed for modern Islamic mobile and web applications — served via a Cloudflare Edge CDN with native… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data.audioautomatic-speech-recognition10K<n<100K5 likes22k downloads1d agoHugging Face02davidscripka /MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html. The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation. The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.audioaudio-classificationn<1K9 likes16k downloads3y agoHugging Face03NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes11k downloads2mo agoHugging Face04Reverb /voxceleb2 VoxCeleb2 Dataset This is the VoxCeleb2 dataset, a large-scale speaker identification dataset. Dataset Description VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube. Files vox2_dev_mp4_part*: Multipart archive containing MP4 video files vox2_dev_txt: Text files with speaker/utterance metadata vox2_meta.csv: Dataset metadata Usage To extract the multipart archive: # Using 7zip 7z x… See the full description on the dataset page: https://huggingface.co/datasets/Reverb/voxceleb2.automatic-speech-recognition100K<n<1M22 likes7.8k downloads1y agoHugging Face05KBLab /rixvox-v2 RixVox-v2: A Swedish parliamentary speech dataset RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.audioautomatic-speech-recognition1M<n<10M12 likes6.7k downloads1y agoHugging Face06RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes5.9k downloads3mo agoHugging Face07RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes5.6k downloads5mo agoHugging Face08parler-tts /libritts_r_filtered Dataset Card for Filtered LibriTTS-R This is a filtered version of LibriTTS-R. It has been filtered based on two sources: LibriTTS-R paper [1], which lists samples for which speech restoration have failed LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected. LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.audiotext-to-speech100K<n<1M24 likes4.7k downloads2y agoHugging Face09lingamvamshikrishnareddy /ramanv-tts-all-rawgated ramanv-tts-all-raw Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages. textautomatic-speech-recognition1M<n<10M0 likes4.5k downloads11d agoHugging Face10SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes4.3k downloads1y agoHugging Face11Reubencf /fma-labeled FMA Labeled — Multi-Attribute Music Dataset 🏆 Submitted to the Uncharted Data Challenge hosted by Adaption Labs — credit to Adaptive Data by Adaption for organizing the hackathon. A large-scale labeled music dataset built on top of the Creative-Commons subset of the Free Music Archive (FMA). Every track has been automatically annotated with lyrics, genre, mood, instruments, tempo, key, and more using Google Gemini (gemini-flash-latest). Intended for training and evaluating music… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/fma-labeled.audioaudio-classification10K<n<100K5 likes4.3k downloads5mo agoHugging Face12joujiboi /Galgame-VisualNovel-Reupload Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of this reupload is to restructure the data for easier and more efficient use with the datasets library, instead of having to manually extract each archive file and parse json files of the original dataset. Loading the entire dataset To load and stream all voice lines from all games combined, simply load the train split. The… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/Galgame-VisualNovel-Reupload.audioautomatic-speech-recognition1M<n<10M38 likes4.1k downloads1y agoHugging Face13Nanboy /RVCBench RVCBench RVCBench is a benchmark dataset for studying robustness in voice cloning, text-to-speech, speaker privacy, audio protection, adversarial audio perturbations, and related audio generation pipelines. Dataset page: https://huggingface.co/datasets/Nanboy/RVCBench Code repository: https://github.com/Nanboy-Ronan/RVCBench Paper: https://arxiv.org/abs/2602.00443 RVCBench is designed for evaluating how modern voice cloning (VC), TTS, and audio generation systems behave under… See the full description on the dataset page: https://huggingface.co/datasets/Nanboy/RVCBench.audiotext-to-speech10K<n<100K0 likes3.1k downloads2mo agoHugging Face14TaurenMountain /REAL-PS4 REAL-PS4 Overview REAL-PS4 is a training dataset for target speaker extraction (TSE) and overlapping speech recognition in real-world multi-speaker meetings. It is constructed from four public meeting corpora — AISHELL-4, AliMeeting, AMI, and CHiME-6 — following the REAL-T benchmark data preparation pipeline. Each sample consists of: A mixture utterance (overlapping speech clip from a real meeting) An enrollment utterance (clean speech from the target speaker in… See the full description on the dataset page: https://huggingface.co/datasets/TaurenMountain/REAL-PS4.audioautomatic-speech-recognition10K<n<100K6 likes2.9k downloads3mo agoHugging Face15retkowski /ytseg YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.audiotoken-classification100K<n<1M8 likes2.7k downloads2mo agoHugging Face16linagora /SUMM-RENote: if the data viewer is not working, use the "example" subset. SUMM-RE The SUMM-RE dataset is a collection of transcripts of French conversations, aligned with the audio signal. It is a corpus of meeting-style conversations in French created for the purpose of the SUMM-RE project (ANR-20-CE23-0017). The full dataset is described in Hunter et al. (2024): "SUMM-RE: A corpus of French meeting-style conversations". Created by: Recording and manual correction of the corpus was… See the full description on the dataset page: https://huggingface.co/datasets/linagora/SUMM-RE.audioautomatic-speech-recognitionn<1K5 likes2.2k downloads2y agoHugging Face17RidheshBhati /tts_farm Multilingual TTS/ASR Aggregated Dataset Cleaned, deduplicated and loudness-normalized Arabic, Japanese, Korean, Turkish, and Vietnamese speech. The training columns are audio (16-bit PCM WAV, 22050 Hz) and text; the remaining columns contain quality and provenance metadata. audiotext-to-speech100K<n<1M2 likes2.2k downloads2mo agoHugging Face18alvanlii /cantonese-radiogated Cantonese Radio Pseudo-Transcription Dataset Contains 14k hours of audio sourced from Archive.org Columns order_index: Represents the order of the audio compared to those from the same filename link: Link of the original full audio transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.audioautomatic-speech-recognition1M<n<10M25 likes2.2k downloads2y agoHugging Face19Reza2kn /gooshkon-chunks Gooshkon MP3 chunks A single-column Hugging Face Audio dataset. Every row contains embedded, playable MP3 bytes in the audio column. Chunks target approximately 30 seconds and are selected at detected quiet intervals with a 12-48 second safety range. The source recordings are not transcript-aligned; this release uses acoustic silence boundaries to avoid cutting through words whenever a usable pause is available. Dataset runtime The dataset contains approximately 2… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/gooshkon-chunks.audioautomatic-speech-recognition100K<n<1M2 likes2k downloads17d agoHugging Face20langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes1.7k downloads2mo agoHugging Face21RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M4 likes1.7k downloads8h agoHugging Face22its5Q /biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio. audiotext-to-speech100K<n<1M23 likes1.3k downloads1y agoHugging Face23Rosie-Lab /BERSt BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Read the paper here Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19 phone positions 7 emotion classes 3 vocal intensity levels varied regional and non-native English accents nonsense phrases covering all English Phonemes Data collection The BERSt dataset represents data… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/BERSt.audioautomatic-speech-recognition1K<n<10K6 likes1.3k downloads1y agoHugging Face24risaleinur /risale-i-nur-sohbet Risale-i Nur Sohbet Prof. Dr. Şener Dilek’ten izin alındı. Türkçe Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri kümelerine karıştırılmaz. Kapsam 2095 sohbet, 954.66 saat 16 kHz mono FLAC ses Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.audioautomatic-speech-recognition1M<n<10M1 likes1.2k downloads18d agoHugging Face25vnahata /OmnilingualASR-retrieval Omnilingual ASR speech-text retrieval (MTEB) Read speech paired with its human transcription, for languages that no existing MTEB audio task covers. Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated transcripts are dropped, since one would otherwise be relevant to several recordings while only one is marked correct. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.audioautomatic-speech-recognition1K<n<10K0 likes1.2k downloads23d agoHugging Face26RobotsMali /afvoices 📘 African Next Voices – Bambara (AfVoices) The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversational settings and annotated using a semi-automated transcription pipeline combining ASR pre-labels and human corrections. We release all the data processing code on GitHub.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices.audioautomatic-speech-recognition100K<n<1M7 likes1.1k downloads21d agoHugging Face27TheNHz /ellipsis-lrs3-rawgated LRS3-TED — verified mirror A mirror of the LRS3-TED dataset (Lip Reading Sentences 3), preserved because the official distribution has been discontinued. This repository adds no new data: it is a re-hosted copy with a full verification report against the official file list, so you know exactly what is and is not here. Attribution LRS3-TED was created by Triantafyllos Afouras, Joon Son Chung and Andrew Zisserman (Visual Geometry Group, University of Oxford): T.… See the full description on the dataset page: https://huggingface.co/datasets/TheNHz/ellipsis-lrs3-raw.textautomatic-speech-recognition1K<n<10K8 likes1k downloads2mo agoHugging Face28rabah2026 /Quran-Ayah-Corpus Quran-Ayah-Corpus: A Multi-Reciter Arabic Quranic Speech Dataset Dataset Description: Ayah-Corpus is a large-scale, multi-reciter Arabic speech dataset meticulously curated for Automatic Speech Recognition (ASR) tasks. It consists of high-quality audio recordings of Quranic verses (Ayahs) paired with their corresponding exact transcriptions. The audio is sourced from two primary repositories: Al-Quran.cloud and EveryAyah.com. This dataset is specifically designed to… See the full description on the dataset page: https://huggingface.co/datasets/rabah2026/Quran-Ayah-Corpus.audioautomatic-speech-recognition100K<n<1M2 likes956 downloads1y agoHugging Face29carlosdanielhernandezmena /ravnursson_asr Dataset Card for ravnursson_asr Dataset Summary The corpus "RAVNURSSON FAROESE SPEECH AND TRANSCRIPTS" (or RAVNURSSON Corpus for short) is a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications in the language that is spoken at the Faroe Islands (Faroese). It was curated at the Reykjavík University (RU) in 2022. The RAVNURSSON Corpus is an extract of the "Basic Language Resource Kit 1.0" (BLARK 1.0) [1] developed… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/ravnursson_asr.audioautomatic-speech-recognition10K<n<100K3 likes955 downloads1y agoHugging Face30RodolfoRizzi /MyMentorLLM-dataset Dataset Card for MyMentorLLM This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information). Dataset Summary MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.audiotext-generation10K<n<100K0 likes946 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.