CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01H-Liu1997 /BEAT2audio1K<n<10K11 likes23k downloads3y agoHugging Face02plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes4.1k downloads9mo agoHugging Face03m-hamza-mughal /beat2-additional-annotations BEAT2 Official Release + Additional Annotations This is a fork of H-Liu1997/BEAT2 that adds annotations contributed by the RAG-Gesture (CVPR 2025) and MIBURI (CVPR 2026) projects. The base BEAT2-English data (motion, audio, TextGrids, semantic labels, pretrained motion-autoencoder weights) is inherited verbatim from upstream; the additional annotations from RAG-Gesture and MIBURI are pushed on top. Citations If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.audio1K<n<10K0 likes2.6k downloads3mo agoHugging Face04committa /serena-synthetic-it-28h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.audiotext-to-speech10K<n<100K1 likes417 downloads2mo agoHugging Face05labhamlet /TUT2018-ov2audio10K<n<100K0 likes225 downloads8mo agoHugging Face06plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes211 downloads23d agoHugging Face07ggfox00000 /dia-earning21-all Earnings 21 The Earnings 21 dataset ( also referred to as earnings21 ) is a 39-hour corpus of earnings calls containing entity dense speech from nine different financial sectors. This corpus is intended to benchmark automatic speech recognition (ASR) systems in the wild with special attention towards named entity recognition (NER). This work has been recently accepted to Interspeech 2021! File Format Overview In the following section, we provide an overview of the file… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-earning21-all.audion<1K0 likes197 downloads5mo agoHugging Face08khaledalganem /sada2022 Dataset Card for SADA صدى Dataset Summary يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر. ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.audio100K<n<1M4 likes188 downloads2y agoHugging Face09outlawmold /sinhala-tts-dataset-archive-20260429-082457 Sinhala TTS Dataset Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare. Stats Metric Value Utterances 218 Train 208 Val 10 Hours 0.51 Mean duration 8.5s Sample rate 22050 Hz Pipeline Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 -> Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB) Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.audiotext-to-speechn<1K0 likes163 downloads5mo agoHugging Face10labhamlet /TUT2018-ov3Copyright (c) 2018 Tampere University of Technology and its licensors All rights reserved. Permission is hereby granted, without written agreement and without license or royalty fees, to use and copy the TUT Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset (“Work”) described in this document and composed of audio and metadata. This grant is only for experimental and non-commercial purposes, provided that the copyright notice in its entirety appear in all… See the full description on the dataset page: https://huggingface.co/datasets/labhamlet/TUT2018-ov3.audio10K<n<100K0 likes155 downloads8mo agoHugging Face11labhamlet /TUT2018-ov1audio1K<n<10K0 likes154 downloads8mo agoHugging Face12MahiA /TUT2017 TUT2017 This is an audio classification dataset for Acoustic Scene Classification. Classes = 15   ,   Split = four-fold Structure audios folder contains audio files. csv_files folder contains CSV files for four-fold cross-validation. To perform cross-validation on fold 1, train_1.csv will be used for the training split and test_1.csv for the testing split, with the same pattern followed for the other folds. To perform training and testing witout cross-validation, use… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/TUT2017.audio10K<n<100K0 likes134 downloads2y agoHugging Face13committa /serena-synthetic-it-27h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.audiotext-to-speech10K<n<100K1 likes121 downloads2mo agoHugging Face14aalshalfi /sada2022-arabic-tts SADA 2022 - Saudi Arabic Dataset for TTS مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech) المصدر الأصلي Kaggle: sdaiancai/sada2022 الاستخدام # طريقة 1: Git Clone !git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset # طريقة 2: مكتبة datasets from datasets import load_dataset dataset = load_dataset("aalshalfi/sada2022-arabic-tts") الملفات valid.csv - ملف البيانات الرئيسي wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.audio100K<n<1M0 likes117 downloads8mo agoHugging Face15ag2003 /bhavvani Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning This repository contains the BhavVani dataset introduced in the INTERSPEECH 2024 Paper : Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning Please fill this form for accessing the audio files associated with the BhavVani dataset: Form Link Overview In our work, we propose the following contributions:… See the full description on the dataset page: https://huggingface.co/datasets/ag2003/bhavvani.textautomatic-speech-recognition1K<n<10K2 likes107 downloads5mo agoHugging Face16Sundus246 /SADA_khaledalganemsada2022_Rawdate Dataset Card for SADA صدى Dataset Summary يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر. ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر… See the full description on the dataset page: https://huggingface.co/datasets/Sundus246/SADA_khaledalganemsada2022_Rawdate.audio100K<n<1M0 likes102 downloads2mo agoHugging Face17Wonder239 /DEAF DEAF DEAF is a collection of audio data, aligned text metadata, and data-generation scripts accompanying the paper DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models. This repository is organized as a Hugging Face dataset repository and contains the locally hosted resources used in the paper: BSC audio, SIC audio, paired text metadata, and the scripts used to generate the speech-related subsets. Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/Wonder239/DEAF.audioaudio-classificationn<1K0 likes80 downloads3mo agoHugging Face18labhamlet /TUT2018-ov2-SELDaudio10K<n<100K0 likes75 downloads8mo agoHugging Face19mueller91 /human-perception-audio-deepfake-2026 Human Audio Deepfake Perception 2026 A large-scale listening study evaluating how well humans detect modern audio deepfakes. The dataset contains 35,532 deepfake-detection judgments from 1,768 anonymous participants across 138 TTS and voice-conversion systems, collected via a publicly accessible online listening game in 2025–2026. This is the successor to the 2021 ASVspoof-2019 perception study (Müller, Pizzi & Williams, 2022) and extends the same paradigm to modern systems… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/human-perception-audio-deepfake-2026.textaudio-classification10K<n<100K4 likes70 downloads4mo agoHugging Face20Abdelrahman2922 /arabic-tts-saudi-multi-speaker-xtts Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦 This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format. 📌 Overview Language: Arabic (Saudi Dialect) Format: LJSpeech Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.) Speakers: Multi-speaker (Male & Female) Audio Format: WAV (mono recommended) Sample Rate: 22050 Hz (recommended) 📂 Structure all_data/ │ ├── wavs/ │ ├── sample_0.wav │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.audio1K<n<10K2 likes68 downloads6mo agoHugging Face21YDX07 /YDX07_Multilingual_Corpus_2026audion<1K0 likes48 downloads26d agoHugging Face22swc2 /aatc2025simaudion<1K0 likes45 downloads1y agoHugging Face23pervasiveaidataresearchlab2025 /VN-SpeechMix_Datasetgated VN-SpeechMix: A Large-Scale Multi-Dialect Vietnamese Speech Mixture Dataset VN-SpeechMix is a large-scale, multi-dialect Vietnamese speech mixture dataset for two-speaker speech separation research. It is built from the ViMD corpus (Van Dinh et al., EMNLP 2024) using a loudness-aware mixing pipeline (LUFS normalization + two-stage anti-clipping) and a dialect-aware pairing strategy across Vietnam's three macro-dialect regions (North / Central / South). 26,000 two-speaker… See the full description on the dataset page: https://huggingface.co/datasets/pervasiveaidataresearchlab2025/VN-SpeechMix_Dataset.tabularaudio-to-audio10K<n<100K0 likes44 downloads19d agoHugging Face24LeyuCompetition /Yadonay-YDX07_Multilingual_Corpus_2026audion<1K0 likes44 downloads22d agoHugging Face25ofc-its-phyla /amharic-speech-dataset-2026 Amharic Speech Dataset 2026 Overview This dataset contains Amharic speech recordings collected using the Leyu Platform for the Leyu Platform Competition 2026. Language Amharic (am) Dialect Standard Addis Ababa Amharic Speaker Information Number of Speakers: 1 Speaker IDs: SPK001 Audio Format Format: M4A Duration: 10–60 seconds per recording Directory Structure audio/ metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/ofc-its-phyla/amharic-speech-dataset-2026.audion<1K0 likes41 downloads2mo agoHugging Face26ShiroOnigami23 /emotion-voice-dataset emotion voice dataset Developed by Aryan Singh Chandel (Shiro) at Rustamji Institute of Technology (RJIT). 📝 Overview This repository contains assets for emotion voice dataset. It is a professional research component of the Shiro AI ecosystem. 🚀 Status The core files are live. Detailed usage instructions and technical benchmarks are currently being compiled for the elite release. audio1K<n<10K5 likes38 downloads9mo agoHugging Face27uam-wmi-asr-eval-labs /2026-dwesui-g01-neurologia DWESUI 2026 - Grupa 1 - neurologia Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny. Zespol (atrybucja): Grupa 1 (DWESUI 2026) Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia Domena: neurologia Licencja zrodla: nagrania YouTube CC-BY + synteza TTS Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.audioautomatic-speech-recognitionn<1K0 likes34 downloads1mo agoHugging Face28TransferRapid /CommonVoices20_ro Common Voices Corpus 20.0 (Romanian) Common Voices is an open-source dataset of speech recordings created by Mozilla to improve speech recognition technologies. It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide. Challenges: The raw dataset included numerous recordings with incorrect transcriptions or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.audioautomatic-speech-recognition10K<n<100K4 likes32 downloads2y agoHugging Face29yangwang825 /vox2-veri-full VoxCeleb 2 VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube. Verification Split train validation test # of speakers 5,994 5,994 118 # of samples 982,808 109,201 36,237 Data Fields ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>. duration (float64): The duration of the segment in seconds. wav (string): The filepath of the waveform. start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-full.tabularaudio-classification1M<n<10M0 likes31 downloads3y agoHugging Face30s479246 /dwesui-grupa-2-kulinarna G2-Polish-Culinary-ASR-Evaluation-Corpus Korpus do ewaluacji systemow ASR jezyka polskiego (domena kulinarna) stworzony w ramach warsztatow Ewaluacja Systemow Rozpoznawania Mowy (UAM WMI, edycja 2026, zespol 2). Publikowany podzbior to mowa naturalna z wideo kulinarnych YouTube (licencja CC-BY) - sluzy do badania odpornosci ASR na szum kuchenny oraz dopasowania domenowego do specjalistycznego slownictwa (zapozyczenia, miary, liczby). Pelny eksperyment ewaluacyjny zespolu… See the full description on the dataset page: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna.audioautomatic-speech-recognitionn<1K0 likes28 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.