CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01meganwei /syntheory Dataset Card for SynTheory Dataset Summary SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions). Each of these 7 concepts has its own config. tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time). time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.audioaudio-classification100K<n<1M13 likes3.5k downloads2y agoHugging Face02PoojasreeBalasubramanian /synthetic-wakewordsaudio10K<n<100K0 likes3.1k downloads2mo agoHugging Face03BertilBraun /voice-light-synthetic-audio Voice-Light Synthetic Audio English-only synthetic conversational speech for training and evaluating streaming turn-taking models. The corpus focuses on end-of-turn prediction, continuation holds, short backchannels, interruptions, and response timing. The dataset contains user-side FLAC speech units plus typed conversation plans, rendering provenance, quality ledgers, and deterministic reconstruction metadata. Assistant speech is represented as a time-varying… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-synthetic-audio.audio10K<n<100K0 likes1.9k downloads29d agoHugging Face04projecte-aina /synthetic_dem Dataset Card for synthetic_dem Dataset Summary The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC). It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.audioautomatic-speech-recognition100K<n<1M2 likes1.5k downloads1y agoHugging Face05laion /synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. https://huggingface.co/datasets/sleeping-ai/Vocal-burst We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories. It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts. audio100K<n<1M6 likes1.5k downloads2y agoHugging Face06Codec-SUPERB /fluent_speech_commands_synth Dataset Card for "fluent_speech_commands_synth" More Information needed audio100K<n<1M1 likes1.3k downloads3y agoHugging Face07Scicom-intl /Synthetic-User-Turn-TTS Synthetic Malaysian Telco Call-Centre Speech Synthetic Malaysian call-centre customer utterances, as text and as speech. The text is fully synthetic dialogue styled after real Malaysian ISP/telco ("Unifi") call-centre recordings, containing no real customer data. The audio subsets take customer (user) turns and voice them with a voice-conversion model, keeping only clips an ASR round-trip confirms are accurate. Subsets subset rows content default 4,260… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Synthetic-User-Turn-TTS.audio100K<n<1M0 likes1k downloads2d agoHugging Face08Silasimo /SynthGT SynthGT A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment Authors Silas Antonisen, Iván López-Espejo Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing. Overview SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations. The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.audioautomatic-speech-recognition1K<n<10K1 likes1k downloads2mo agoHugging Face09Codec-SUPERB /librispeech_synth Dataset Card for "librispeech_synth" More Information needed audio1M<n<10M1 likes929 downloads3y agoHugging Face10synthbot /pony-speechaudiotext-to-speech10K<n<100K20 likes827 downloads2y agoHugging Face11Aalto-Speech-Synthesis /icelandic_asr Icelandic ASR Collection This repository collects six Icelandic speech corpora in directly loadable Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a convenience repackaging: the linked CLARIN-IS records and original dataset repositories remain the canonical sources and should be cited when using the data. No configuration is selected by default. Choose a corpus configuration and, for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.audioautomatic-speech-recognition1M<n<10M0 likes631 downloads22d agoHugging Face12yuriyvnv /capes_synthetic_audio_filteredaudio10K<n<100K0 likes630 downloads1y agoHugging Face13Codec-SUPERB /vocal_imitation_synth Dataset Card for "vocal_imitation_synth" More Information needed audio10K<n<100K1 likes621 downloads3y agoHugging Face14Codec-SUPERB /crema_d_synth Dataset Card for "crema_d_synth" More Information needed audio100K<n<1M0 likes599 downloads3y agoHugging Face15Codec-SUPERB /voxceleb1_synthaudio10K<n<100K3 likes599 downloads3y agoHugging Face16Codec-SUPERB /maestro_synth Dataset Card for "maestro_synth" More Information needed audio1K<n<10K0 likes571 downloads3y agoHugging Face17CodecSR /librispeech_asr_test_48k_synthaudio100K<n<1M0 likes563 downloads3y agoHugging Face18CodecSR /torgo_synthaudio100K<n<1M0 likes558 downloads2y agoHugging Face19CodecSR /vox_lingua_top10_synthaudio10K<n<100K0 likes553 downloads3y agoHugging Face20derekxkwan /syntheory_plus Notes Dataset Viewer is disabled as we wanted to keep everything in WAV format with CSV metadata instead of Parquet files and the dataset is too large for Dataset Viewer to index properly. Dataset Authors Derek Kwan and Patrick Donnelly Related Paper The paper that introduces this dataset is "Probing for Advanced Music Theory Concepts in Generative Music Models" by Derek Kwan and Patrick Donnelly presented at EvoMUSART 2026 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/derekxkwan/syntheory_plus.audio100K<n<1M0 likes546 downloads19d agoHugging Face21CodecSR /vocalset_synthaudio10K<n<100K0 likes545 downloads3y agoHugging Face22CodecSR /speech_accent_archive_synthaudio10K<n<100K0 likes535 downloads2y agoHugging Face23CodecSR /librispeech_asr_test_synthaudio100K<n<1M0 likes507 downloads3y agoHugging Face24Codec-SUPERB /vocalset_synth Dataset Card for "vocalset_synth" More Information needed audio10K<n<100K0 likes498 downloads3y agoHugging Face25Codec-SUPERB /opensinger_synthaudio10K<n<100K0 likes450 downloads3y agoHugging Face26committa /serena-synthetic-it-28h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.audiotext-to-speech10K<n<100K1 likes417 downloads2mo agoHugging Face27CodecSR /voxceleb1_synthaudio100K<n<1M0 likes416 downloads3y agoHugging Face28TigreGotico /synthetic-wakeword-hey_computer synthetic-wakeword-hey_computer Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey computer". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.audioaudio-classification1K<n<10K0 likes406 downloads8d agoHugging Face29bc7ec356 /synthetic-speech-indicaudio100K<n<1M1 likes399 downloads5mo agoHugging Face30CodecSR /easycall_synthaudio100K<n<1M0 likes394 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.