CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01its5Q /biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio. audiotext-to-speech100K<n<1M23 likes1.2k downloads1y agoHugging Face02anyspeech /ipapack_plus_train_3audio1M<n<10M0 likes1k downloads1y agoHugging Face03IWSLT /IWSLT.OfflineTaskaudiotranslationn<1K2 likes227 downloads3y agoHugging Face04itzune /antton-dataset Antton Dataset (Synthetic) This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Antton" voice model. This dataset was generated by Itzune and serves as the primary source for training the itzune/antton-tts (Piper version) model. Dataset Structure Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/antton-dataset.audiotext-to-speech100K<n<1M0 likes150 downloads6mo agoHugging Face05vikrant-vikram /INDICA INDICA: An Audio Indic-Language Telecom Fraud Analysis Benchmark Multilingual | Audio + Text | Benchmark for Fraud Detection Overview INDICA is a comprehensive benchmark for telecom fraud call analysis in Indic languages.It is built on the IndiF dataset, the first large-scale multilingual dataset for fraud detection in telecom conversations. This benchmark enables research in: Scenario Classification Fraud Call Detection Fraud-Type Classification… See the full description on the dataset page: https://huggingface.co/datasets/vikrant-vikram/INDICA.audio1M<n<10M2 likes128 downloads6mo agoHugging Face06issai /Multilingual_Speech_Dataset Multilingual Speech Dataset Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English Repository: https://github.com/IS2AI/MultilingualASR Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.audioautomatic-speech-recognition100K<n<1M3 likes110 downloads2y agoHugging Face07anyspeech /ipapack_plus_5audio1M<n<10M0 likes100 downloads2y agoHugging Face08anyspeech /ipapack_plus_3audio100K<n<1M0 likes58 downloads2y agoHugging Face09anyspeech /ipapack_plus_6audio1M<n<10M0 likes53 downloads2y agoHugging Face10its5Q /bigger-ru-bookaudio10K<n<100K13 likes48 downloads1y agoHugging Face11anyspeech /ipapack_plus_1audio100K<n<1M0 likes44 downloads2y agoHugging Face12Exgc /iter_1audio100K<n<1M0 likes34 downloads2y agoHugging Face13McCheng /IMDAaudio1M<n<10M0 likes31 downloads2y agoHugging Face14itzune /maider-dataset Maider Dataset (Synthetic) This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Maider" voice model. This dataset was generated by Itzune and serves as the primary source for training the itzune/maider-tts (Piper version) model. Dataset Structure Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/maider-dataset.audiotext-to-speech10K<n<100K0 likes31 downloads7mo agoHugging Face15anyspeech /ipapack_plus_4audio1M<n<10M0 likes30 downloads2y agoHugging Face16kehanlu /Speech-IFEvalaudio1K<n<10K0 likes30 downloads2y agoHugging Face17ivanj-0 /audio-qformeraudio1M<n<10M0 likes28 downloads1y agoHugging Face18VoiceNet /improved-synthetic-vocal-burtsaudio10K<n<100K0 likes22 downloads5mo agoHugging Face19pengyizhou /IALP-2026-data IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC Adaptation." This repository holds the fixed target-query, validation, and evaluation sets used across all experiments. Each part is a self-contained .tar.gz. All audio is 16 kHz mono. Each split ships with: audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16) wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.audioautomatic-speech-recognition10K<n<100K0 likes21 downloads3mo agoHugging Face20laion /improved_synthetic_vocal_burtsaudio10K<n<100K4 likes16 downloads11mo agoHugging Face21syn-omni-sony /internvideo2_dataaudio100K<n<1M0 likes16 downloads6mo agoHugging Face22CJY /Chinese-Dialogue-180k-Instruct-Audioaudio100K<n<1M4 likes13 downloads1y agoHugging Face23VoiceNet /multilingual-in-the-wildaudio100K<n<1M0 likes13 downloads5mo agoHugging Face24sleeping-ai /ImagineV1-audioaudio1K<n<10K0 likes12 downloads1y agoHugging Face25subahponraj /irescvttaudion<1K0 likes11 downloads2y agoHugging Face26HKUSTAudio /AudioX-IFcapsgated [ICLR 2026] AudioX-IFcaps: Instruction-Following Audio Caption Dataset AudioX-IFcaps (Instruction-Following) is a large-scale, high-quality multimodal dataset designed for training unified audio and music generation models. The dataset contains over 7 million samples with fine-grained, structured annotations that enable precise control over audio generation, including sound event categories, counts, temporal ordering, and timestamps. 📊 Dataset Statistics General Audio:… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/AudioX-IFcaps.audiotext-to-audio100K<n<1M7 likes11 downloads6mo agoHugging Face27dangkhoadl /ICASSP2024-Acoustic_Scattering_AI-Noninvasive_Object_Classificationsaudio10K<n<100K0 likes8 downloads3y agoHugging Face28marcosremar2 /pt-br-tts-iasmin-qwen3 pt-br-tts-iasmin-qwen3 21957 clips PT-BR sintetizados com Qwen3-TTS-12Hz-1.7B. Voz Iasmin (voice-clone, is_iasmin=true, ~13957 clips) + vozes diversas CustomVoice (Ryan, Aiden, Vivian, Dylan, is_iasmin=false). WAV em tar shards (WebDataset); transcricao, voz, is_iasmin e sr em metadata.jsonl. audiotext-to-speech10K<n<100K0 likes8 downloads4mo agoHugging Face29IamXiangyu /osu-beatmaps-duplicated osu! Beatmaps Dataset (WebDataset) A collection of ranked/loved osu! beatmaps with audio and chart data, in WebDataset format. Dataset Variants Variant Audio Format Description original MP3/OGG/WAV Full quality original audio files compressed 64kbps Mono Opus Compressed audio for smaller download from datasets import load_dataset # Load original audio variant ds = load_dataset("project-riz/osu-beatmaps", "original", streaming=True) # Load compressed… See the full description on the dataset page: https://huggingface.co/datasets/IamXiangyu/osu-beatmaps-duplicated.audioaudio-classification10K<n<100K0 likes7 downloads6mo agoHugging Face30omniway /Audio_speaker_needle_in_haystackaudio1K<n<10K1 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.