CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /seamless-interaction Seamless Interaction Dataset A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research 🖼️ Blog 🌐 Website 🎮 Demo 📦 GitHub 📄 Paper Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in… See the full description on the dataset page: https://huggingface.co/datasets/facebook/seamless-interaction.audio198 likes107k downloads1y agoHugging Face02hf-internal-testing /audiofolder_single_config_in_metadataaudion<1K0 likes102k downloads3y agoHugging Face03artur-muratov /multilingual-speech-commands-15lang Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.audio1M<n<10M16 likes80k downloads1y agoHugging Face04japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes46k downloads2y agoHugging Face05SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes40k downloads1y agoHugging Face06MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M285 likes37k downloads2y agoHugging Face07sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face08MERaLiON /Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching. ASR: Automatic Speech Recognition SQA: Speech Question Answering SDS: Spoken Dialogue Summarization PQA: Paralinguistic Question Answering from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.audio10M<n<100M22 likes28k downloads2y agoHugging Face09MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face10VoiceHub /voicehub-arena-seed-tts-eval VoiceHub Arena — native TTS evaluations Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores. Interactive demo · Source code Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.audiotext-to-speech0 likes25k downloads4d agoHugging Face11arsaporta /symile-m3 Dataset Card for Symile-M3 Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice. Paper: https://arxiv.org/abs/2411.01053 GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.audiozero-shot-classification10M<n<100M8 likes23k downloads2y agoHugging Face12S3Sound /acidaudio1K<n<10K1 likes23k downloads1y agoHugging Face13MLCommons /speech-wikimedia Dataset Card for Speech Wikimedia Dataset Summary The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers. Each audiofile should have one or more transcriptions in different languages. Transcription languages English German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.audion<1K14 likes23k downloads3y agoHugging Face14cksqs /SynStard-1000 SynStard-1000 Dataset Summary SynStard-1000 is a 1,000-hour synthetic dataset for training and evaluating end-to-end speech-to-speech translation (S2ST) models. It is built from English-Chinese parallel texts in the WMT News Commentary v18 corpus and contains approximately 390,000 sentence pairs with paired synthetic speech. Dataset Structure . ├── map/ │ └── all.tsv │── text/ │ ├── en/ │ │ ├── en.txt │ │ ├── en_1.txt │ │ ├── ... │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/cksqs/SynStard-1000.audio100K<n<1M2 likes22k downloads10mo agoHugging Face15speechcolab /gigaspeechgated Dataset Card for Gigaspeech Dataset Description GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. Example Usage The training split has several configurations of… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech.audioautomatic-speech-recognition10M<n<100M173 likes18k downloads8mo agoHugging Face16Sinoosoida /SpeechRu Russian Podcasts (unlabeled) ~186k unlabeled Russian-language podcast episodes scraped from the web, packaged as Parquet shards with the audio bytes embedded. The audio has no transcripts — this is an unsupervised / self-supervised audio corpus, suitable for ASR pre-training, speech-representation learning, TTS data mining, audio classification, and similar tasks. Each row contains: audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo), decoded on-the-fly via the… See the full description on the dataset page: https://huggingface.co/datasets/Sinoosoida/SpeechRu.audioautomatic-speech-recognition100K<n<1M4 likes17k downloads3mo agoHugging Face17RidheshBhati /Complete_Data_Source_100K_HOURS Multi-Language Audio Collection (100K Hours) This repository is physically reorganized for Absolute 100% Data Visibility. 🏗️ Global Consolidator Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here. audio1M<n<10M4 likes16k downloads5mo agoHugging Face18laion /soundscapesaudio10M<n<100M7 likes16k downloads1y agoHugging Face19simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes15k downloads23d agoHugging Face20EwanB /satb-choral-dataset SATB Choral Source Separation Dataset (Compressed) This dataset contains preprocessed 4-second audio chunks from the Choral Singing Dataset (CSD) formatted for SATB (Soprano, Alto, Tenor, Bass) voice source separation tasks. This is the compressed version with 8kHz sample rate and int8 precision for smaller file sizes. Data Structure Folder Structure ├── chunks/ # All individual chunk .pt files ├── quality_samples/ # Sample WAV… See the full description on the dataset page: https://huggingface.co/datasets/EwanB/satb-choral-dataset.audioaudio-classification1K<n<10K0 likes15k downloads1y agoHugging Face21GuyHam /c-sac-corpora C-SAC LibriTTS-R training subset Deterministically selected and resampled speech from mythicinfinity/libritts_r for the C-SAC causal speech-codec program. The package retains source revision, Parquet shard, row, utterance, transcript, and content hashes. LibriTTS-R is distributed under CC BY 4.0; downstream users remain responsible for attribution. Only prefixes with a hash-bound _COMPLETE.json sentinel are admissible. audiotext-to-speech100K<n<1M0 likes14k downloads1d agoHugging Face22shenyunhang /AISHELL-4 AISHELL-4 Identifier: SLR111 Summary: A Free Mandarin Multi-channel Meeting Speech Corpus, provided by Beijing Shell Shell Technology Co.,Ltd Category: Speech License: CC BY-SA 4.0 Downloads (use a mirror closer to you): train_L.tar.gz [7.0G] &nbsp; ( Training set of large room, 8-channel microphone array speech ) &nbsp; Mirrors: [US] &nbsp; [EU] &nbsp; [CN] &nbsp; train_M.tar.gz [25G] &nbsp; ( Training set of medium room, 8-channel microphone array speech ) &nbsp;… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-4.audio10K<n<100K2 likes14k downloads2y agoHugging Face23mshah1 /speech_robust_bench Dataset Card for "speech_robust_bench" More Information needed audio1M<n<10M11 likes14k downloads2y agoHugging Face24hf-internal-testing /dummy-audio-samplesaudion<1K0 likes13k downloads2d agoHugging Face25ilanashapiro /stg-paired-audioaudio0 likes11k downloads1y agoHugging Face26juneys /shamupiaudion<1K0 likes11k downloads1mo agoHugging Face27Scicom-intl /semantic-vad-eot Semantic-VAD EOT End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible with livekit/eot-bench-data. Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered silence_spans. Per the eot-bench convention the last silence span is the true end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored). Splits For every data type, all shards except the last form the train base; that… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/semantic-vad-eot.audiovoice-activity-detection10M<n<100M11 likes11k downloads5d agoHugging Face28sanganaka /Vedavani-Dataset Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry Vedavani is the first benchmark dataset for automatic speech recognition (ASR) on Vedic Sanskrit poetry, consisting of richly annotated verses from the Rig Veda and Atharva Veda. This corpus captures the unique prosodic structure, phonetic complexity, and chanting style found in traditional Vedic recitation. 🔗 Paper: Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry (ACL 2025)📁 GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/Vedavani-Dataset.audioautomatic-speech-recognition0 likes10k downloads1y agoHugging Face29DeepSeekOracle /excavationpro-music-stream Excavationpro public music stream (160 kbps) Owner / artist: Justin Helmer · Excavationpro · LightfatherPolicy: Own-work only. Public discovery streams (not DistroKid-dependent).Lattice signature: Δ9Φ963-PUBLIC-MUSIC-STREAM-v1 Listen https://deepseekoracle.github.io/Excavationpro/excavationpro-listen.html http://asiancoastline.com/ (custom domain music portal) Layout Path Role stream/<sha256>.mp3 Flat 160k streams (~first 10k −… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/excavationpro-music-stream.audioaudio-to-audio10K<n<100K0 likes10k downloads20h agoHugging Face30sarulab-speech /mls_sidon MLS-Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. The dataset is provided in WebDataset format for efficient large-scale training. Source: Multilingual LibriSpeech Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese Format: WebDataset (.tar shards) License: CC-BY-4.0 Dataset Structure Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.audiotext-to-speech10M<n<100M11 likes9.8k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.