CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DreamyWanderer /Slakh2100-FLAC-Redux-Reducedaudion<1K0 likes1.5k downloads3y agoHugging Face02shb777 /gemini-flash-2.0-speech 🎙️ Gemini Flash 2.0 Speech Dataset This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English. 🏅 #1 Trending Audio Dataset in Feb 2025 🏅 Used in training of Kokoro TTS and LLaSA 1B 〽️ Stats Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours) Average duration: 10.83 seconds Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.audiotext-to-speech10K<n<100K60 likes1.3k downloads1y agoHugging Face03HKUSTAudio /Audio-FLAN-Datasetgated Audio-FLAN Dataset (Paper) (the FULL audio files and jsonl files are still updating) An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound. 1. Dataset Structure The Audio-FLAN-Dataset has the following directory structure: Audio-FLAN-Dataset/ ├── audio_files/ │ ├── audio/ │ │ └── 177_TAU_Urban_Acoustic_Scenes_2022/ │ │ └── 179_Audioset_for_Audio_Inpainting/ │ │ └── ... │ ├── music/ │ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.audiotext-to-speech10M<n<100M47 likes1.2k downloads1y agoHugging Face04roro128 /musdb18-hq-flac MUSDB18-HQ (FLAC Optimized) Only the audio payload is converted to lossless PCM-16 FLAC. The original columns are preserved: audio, path, and instrument. Source dataset This dataset is derived from the original MUSDB18-HQ dataset. The original dataset card and license are the authoritative references for the source audio and annotations. Only the audio payload was transcoded to lossless PCM-16 FLAC; paths, instrument labels, and source track structure were… See the full description on the dataset page: https://huggingface.co/datasets/roro128/musdb18-hq-flac.audion<1K0 likes680 downloads2mo agoHugging Face05nccratliri /wing-flap-noise-audio-examplesaudion<1K0 likes679 downloads2y agoHugging Face06jzju /wavenet_flashback Dataset Card for "wavenet_flashback" https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig sv-SE-Wavenet-{voice} https://spraakbanken.gu.se/resurser/flashback-dator audioautomatic-speech-recognition10K<n<100K0 likes654 downloads4y agoHugging Face07DreamyWanderer /MAESTRO-v3.0-FLACaudion<1K0 likes477 downloads3y agoHugging Face08hf-internal-testing /dummy-flac-single-exampleaudion<1K0 likes378 downloads4y agoHugging Face09badrex /ethiopian-speech-flat Ethio Speech Copus — Afrivoices Ethiopian 📌 Overview The Ethio Speech Corpus dataset is a multilingual speech corpus containing audio–text pairs across five Ethiopian languages. It is designed to support the development of speech-to-text technologies for low-resource languages. This dataset is part of the Afrivoices initiative — a collaborative effort to create a large-scale ASR dataset for African languages. The broader goal of the initiative is to collect 600 hours… See the full description on the dataset page: https://huggingface.co/datasets/badrex/ethiopian-speech-flat.audioautomatic-speech-recognition100K<n<1M2 likes293 downloads8mo agoHugging Face10Flamme-VRM /kazakh-speech-dataset Kazakh Speech Dataset If you find this dataset helpful please press 'like' button Summary The Kazakh Speech Dataset is a large-scale open-source speech corpus for the Kazakh language. This dataset is designed to support the development of Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems for the Kazakh language. Dataset Statistics Total Audio Duration: ~726 hours Language: Kazakh (kk) Audio Format: FLAC Sampling Rate: 16kHz Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Flamme-VRM/kazakh-speech-dataset.audio100K<n<1M3 likes278 downloads9mo agoHugging Face11Reza2kn /Wikipedia-FA-EN-DeepSeek-V4-Flash-0731 Wikipedia Persian to English — DeepSeek V4 Flash 0731 Rolling, machine-generated English translations of Persian Wikipedia articles from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration full_articles_fa_without_en. 129,816 translations are currently published in 26 immutable Parquet shards. The target release contains 129,816 translations; five source rows have empty plain_text and are not translated. Shards are published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.audiotranslation100K<n<1M1 likes213 downloads27d agoHugging Face12mteb /FLARE-1k-Unified-T2VAaudio1K<n<10K0 likes195 downloads2mo agoHugging Face13fireblade2534 /Gemini-2.0-Flash-Aoede-Voiceaudiotext-to-speech1K<n<10K8 likes193 downloads2y agoHugging Face14mteb /FLARE-1k-Audio-T2VAaudio1K<n<10K0 likes191 downloads2mo agoHugging Face15fireblade2534 /Gemini-2.0-Flash-Fenrir-Voiceaudiotext-to-speech1K<n<10K4 likes179 downloads2y agoHugging Face16fireblade2534 /Gemini-2.0-Flash-Kore-Voiceaudiotext-to-speech1K<n<10K4 likes164 downloads2y agoHugging Face17flagship-ai /ghomala-spoken-bible Ghomálá' Spoken New Testament — aligned audio + trilingual text Part of the Lingo / NativeAI language-preservation project. This is ~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter with parallel text in Ghomálá', French, and English. Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.audioautomatic-speech-recognition10K<n<100K0 likes123 downloads4mo agoHugging Face18Wissam42 /FLARE-1k-Audio-T2VAaudio1K<n<10K0 likes76 downloads2mo agoHugging Face19akuzdeuov /gemini_flash_ttsaudio10K<n<100K0 likes70 downloads1y agoHugging Face20roro128 /fleurs-flac FLEURS-FLAC A losslessly FLAC-compressed version of Google's FLEURS dataset covering 102 languages. Overview This repository contains the Google FLEURS dataset repackaged into Parquet shards with PCM24 FLAC-compressed audio binaries. Key points: Audio streams are converted to FLAC (PCM24) with sample-level PCM verification against the source. Sharded into ~500MB Parquet files per split for efficient I/O and streaming. Covers all 102 languages from the original… See the full description on the dataset page: https://huggingface.co/datasets/roro128/fleurs-flac.tabularautomatic-speech-recognition100K<n<1M0 likes52 downloads2mo agoHugging Face21perigo /flaviaaudion<1K0 likes42 downloads3y agoHugging Face22fireblade2534 /Gemini-2.0-Flash-Puck-Voiceaudiotext-to-speech1K<n<10K3 likes41 downloads2y agoHugging Face23Wissam42 /FLARE-1k-Unified-T2VAaudio1K<n<10K0 likes32 downloads2mo agoHugging Face24windcrossroad /GenderQA-gemini-1.5-flash Dataset Card for "GenderQA-gemini-1.5-flash" More Information needed audion<1K0 likes30 downloads2y agoHugging Face25windcrossroad /EmotionQA-gemini-1.5-flash-fix Dataset Card for "EmotionQA-gemini-1.5-flash-fix" More Information needed audion<1K0 likes29 downloads2y agoHugging Face26windcrossroad /LanguageQA-gemini-1.5-flash Dataset Card for "LanguageQA-gemini-1.5-flash" More Information needed audion<1K0 likes29 downloads2y agoHugging Face27eustlb /audio-arena-audio-flamingo-3-hfaudion<1K0 likes26 downloads6mo agoHugging Face28beastLucifer /music-flamingo-dataset Music Flamingo Distillation Dataset 2025 This dataset contains the Music Flamingo distillation data for training and evaluation. Dataset Description Music Flamingo is a model for music understanding and generation tasks. This dataset was originally hosted on Kaggle and has been migrated to HuggingFace for easier access and integration with the Hugging Face ecosystem. Original Source Originally available at:… See the full description on the dataset page: https://huggingface.co/datasets/beastLucifer/music-flamingo-dataset.audioaudio-to-audion<1K0 likes24 downloads9mo agoHugging Face29pelatihpokemongo /datatalk-flac16kaudio1K<n<10K0 likes23 downloads11mo agoHugging Face30windcrossroad /EmotionQA-gemini-1.5-flash Dataset Card for "EmotionQA-gemini-1.5-flash" More Information needed audion<1K0 likes21 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.