CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hui519 /WildElder WILDELDER: A CHINESE ELDERLY SPEECH DATASET FROM THE WILD WITH FINE-GRAINED MANUAL ANNOTATIONS Paper: https://huggingface.co/papers/2510.09344Code: https://github.com/NKU-HLT/WildElder WildElder is a speech dataset focused on elderly scenarios. It contains raw audio and corresponding text annotations and can be used for ASR, speaker-related tasks, and front-/back-end speech processing research. The data was collected and cleaned from real-world environments to preserve diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Hui519/WildElder.audioautomatic-speech-recognition10K<n<100K3 likes4.4k downloads5mo agoHugging Face02Hezep /AudioMarathon 🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audioaudio-classification1K<n<10K4 likes3.9k downloads11mo agoHugging Face03softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face04BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.4k downloads10mo agoHugging Face05Silasimo /SynthGT SynthGT A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment Authors Silas Antonisen, Iván López-Espejo Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing. Overview SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations. The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.audioautomatic-speech-recognition1K<n<10K1 likes1.3k downloads2mo agoHugging Face06egcortes /asr-jargon-specialized-vocabulary A Dataset for Evaluating ASR on Specialized Vocabulary Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026). Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code Configs Config Language Description synthetic_terms_en English Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms synthetic_terms_pt Portuguese Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes570 downloads2mo agoHugging Face07anonymous-nsc-author /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes319 downloads3mo agoHugging Face08shadow-wxh /VoiceCommandAudioThis is mainly used for fine tune "VoiceCommand" a speech congnition MOD dedicated for SilentHunter game series audioautomatic-speech-recognitionn<1K1 likes252 downloads2y agoHugging Face09paodigitalhub /pao-audio-dataset 🎙️ Pa'O Audio Dataset ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ 📌 Project Summary The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ). Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.audioautomatic-speech-recognitionn<1K1 likes238 downloads10d agoHugging Face10marleen-snsl /sqp-tts-en SQP TTS (English) Synthesized speech for SQPsychConv_qwen-2.5, a synthetic CBT therapist-client dialogue dataset (English). Each configuration below corresponds to one TTS model. Load a single model with: from datasets import load_dataset ds = load_dataset("sinselm/sqp-tts-en", "qwen3-tts") Models included qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512 fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/sqp-tts-en.audiotext-to-speech1K<n<10K0 likes230 downloads3mo agoHugging Face11maleo-ai /maleo-short-1.5H Dataset Card for Maleo Short 1.5H Dataset Description Dataset Summary Maleo Short 1.5H is a manually curated, rigorously annotated speaker diarization dataset designed to benchmark State-of-the-Art (SOTA) models against complex, "in-the-wild" media domains. While modern diarization pipelines excel in controlled acoustic environments (like telephony or reading corpora), they heavily struggle with the overlapping speech, sound effects, and rapid speaker shifts… See the full description on the dataset page: https://huggingface.co/datasets/maleo-ai/maleo-short-1.5H.audioaudio-classificationn<1K3 likes218 downloads4mo agoHugging Face12FBK-MT /fama-data Dataset Description, Collection, and Source The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST) used to train the FAMA models family. The ASR section of FAMA is derived from the MOSEL data collection, including the automatic transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset. The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.tabulartranslation1M<n<10M2 likes196 downloads1y agoHugging Face13MatBar99 /HALAS Dataset Card for HALAS Dataset Summary HALAS (Hallucination Annotations for Large-scale ASR Systems) is a human-annotated dataset of hallucinations produced by modern automatic speech recognition (ASR) systems on real-world speech recordings. The dataset contains span-level hallucination annotations for ASR outputs generated from recordings in the Earnings22 corpus. HALAS was introduced to address a key limitation in prior hallucination research: most existing… See the full description on the dataset page: https://huggingface.co/datasets/MatBar99/HALAS.textautomatic-speech-recognition1K<n<10K0 likes194 downloads3mo agoHugging Face14outlawmold /sinhala-tts-dataset-archive-20260429-082457 Sinhala TTS Dataset Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare. Stats Metric Value Utterances 218 Train 208 Val 10 Hours 0.51 Mean duration 8.5s Sample rate 22050 Hz Pipeline Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 -> Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB) Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.audiotext-to-speechn<1K0 likes166 downloads5mo agoHugging Face15aranemini /central-kurdish-tts4all TTS4All Central Kurdish Speech Dataset Dataset Summary The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish). The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers. The corpus was designed to support: Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.audiotext-to-speech10K<n<100K3 likes147 downloads3mo agoHugging Face16munasco /africanvoices-naija-batch1-summary African Voices Naija Train Metadata Summary This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis. Files included: batch_summary.csv domain_distribution.csv The repository contains metadata summaries only and does not include raw audio. tabularautomatic-speech-recognitionn<1K0 likes147 downloads6mo agoHugging Face17AudioMarathon /AudioMarathon AudioMarathon AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings. The release package in this directory is organized around 11 benchmark tasks spanning meeting summarization, automatic speech recognition, reading comprehension, authenticity detection, music genre classification, acoustic scene classification, emotion recognition, spoken named entity reasoning, sound event detection, speaker gender… See the full description on the dataset page: https://huggingface.co/datasets/AudioMarathon/AudioMarathon.audioaudio-classificationn<1K0 likes144 downloads5mo agoHugging Face18hadou1225 /Hadou-Voice-Dataset Hadou Voice Dataset ハドウ本人が収録した、日本語音声データセットです。 このページで、特徴の異なる2種類のデータセットを公開しています。 配布データ 設定名 内容 音声数 合計時間 v1(おすすめ) Hadou Calm Voice Dataset v1。落ち着いた中音域、AIキャラクター向けボイスが多めの音声データ 966 約114.02分 v0 Hadou ITA Corpus Dataset v1。ITAコーパスを読み上げた自然な話し声 424 約38.95分 v1 には、AICAコーパス500文、ITAコーパス324文、感情・態度付き90文、同文異演技40文、強度段階12文を収録しています。 v1の詳細: v1/README.txt v0の詳細: v0/README.txt 読み込み例 from datasets import load_dataset # 新しい966音声(既定) dataset =… See the full description on the dataset page: https://huggingface.co/datasets/hadou1225/Hadou-Voice-Dataset.audiotext-to-speech1K<n<10K2 likes136 downloads1mo agoHugging Face19arnauquest /original-songs Dataset Card for "original-songs" (Audio + análisis DSP) Dataset Summary Dataset pequeño de canciones originales creadas con IA, cada una con su WAV, letra transcrita automáticamente (Whisper) y un análisis DSP completo (tempo, tonalidad, loudness, features perceptuales) además de detección de contenido explícito. Pensado para quien quiera mejorar modelos open source: extracción de features musicales, clasificación de audio, transcripción y moderación de letras.… See the full description on the dataset page: https://huggingface.co/datasets/arnauquest/original-songs.audioaudio-classificationn<1K1 likes133 downloads1mo agoHugging Face20ivrit-ai /knesset-plenumsgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps. We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts). The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.audioautomatic-speech-recognition1K<n<10K3 likes109 downloads10mo agoHugging Face21ag2003 /bhavvani Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning This repository contains the BhavVani dataset introduced in the INTERSPEECH 2024 Paper : Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning Please fill this form for accessing the audio files associated with the BhavVani dataset: Form Link Overview In our work, we propose the following contributions:… See the full description on the dataset page: https://huggingface.co/datasets/ag2003/bhavvani.textautomatic-speech-recognition1K<n<10K2 likes103 downloads5mo agoHugging Face22Shramadeepd /uyghur-ASR-dataset Uyghur ASR Corpus (Latin Transliteration) A speech corpus for Uyghur automatic speech recognition, with transcriptions in a case-sensitive Latin transliteration scheme. Approximately 23 hours of audio across 9,468 clips. Uyghur is a Turkic language spoken by roughly 10–12 million people. It is severely under-represented in open speech datasets, and this corpus is intended to support ASR research for the language. Dataset summary Language Uyghur (ug)… See the full description on the dataset page: https://huggingface.co/datasets/Shramadeepd/uyghur-ASR-dataset.audioautomatic-speech-recognition1K<n<10K0 likes97 downloads7d agoHugging Face23s512757 /polish-tedx-asr-eval Polish-TEDx-ASR-Eval A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks. Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1. Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.audioautomatic-speech-recognitionn<1K0 likes88 downloads3mo agoHugging Face24UniDataPro /slovenian-speech-recognition Slovenian Speech Dataset Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.audioautomatic-speech-recognitionn<1K2 likes86 downloads1mo agoHugging Face25marleen-snsl /annomi-tts-en AnnoMI TTS (English) Synthesized speech for the AnnoMI motivational interviewing dialogues (English). Each configuration below corresponds to one TTS model. Load a single model with: from datasets import load_dataset ds = load_dataset("sinselm/annomi-tts-en", "qwen3-tts") Models included qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512 fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/annomi-tts-en.audiotext-to-speechn<1K0 likes86 downloads3mo agoHugging Face26kibaraki /Shinekhen-BuryatAudio collected by Yamakoshi (Tokyo University of Foreign Studies), originally uploaded here (CC BY-SA 4.0). start_time and end_time are from the original audio clips; the audio uploaded here are already converted into per-sentence audio clips. Used in [paper] [GitHub] audioautomatic-speech-recognition1K<n<10K0 likes85 downloads1y agoHugging Face27Wonder239 /DEAF DEAF DEAF is a collection of audio data, aligned text metadata, and data-generation scripts accompanying the paper DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models. This repository is organized as a Hugging Face dataset repository and contains the locally hosted resources used in the paper: BSC audio, SIC audio, paired text metadata, and the scripts used to generate the speech-related subsets. Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/Wonder239/DEAF.audioaudio-classificationn<1K0 likes83 downloads3mo agoHugging Face28FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes81 downloads5mo agoHugging Face29nyuuzyou /asmr Dataset Card for ASMR Audio Dataset Dataset Summary This dataset contains a large collection of ASMR (Autonomous Sensory Meridian Response) audio clips with corresponding machine-generated transcriptions. The dataset includes approximately 283,132 audio segments totaling over 307 hours of content, with an average duration of 3.92 seconds per clip. All audio files are provided in WAV format at 24 kHz sampling rate, making them suitable for various audio processing and… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/asmr.textautomatic-speech-recognition100K<n<1M5 likes80 downloads1y agoHugging Face30IbrahimSalah /The_Arabic_News_speech_Corpus_Dataset Arabic News Speech Corpus Dataset This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics. Dataset Details Dataset Description This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.audioautomatic-speech-recognition1K<n<10K6 likes70 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.