CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ylacombe /cml-tts Dataset Card for CML-TTS Dataset Summary CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.audiotext-to-speech1M<n<10M36 likes119k downloads3y agoHugging Face02espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M33 likes77k downloads1y agoHugging Face03sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face04malaysia-ai /malaysian-youtube Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs data at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube/data Notebooks at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube How to load the data efficiently? import pandas as pd import json from datasets import Audio from torch.utils.data import DataLoader, Dataset chunks = 30 sr = 16000 class Train(Dataset):… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-youtube.audio10K<n<100K5 likes12k downloads2y agoHugging Face05NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes11k downloads2mo agoHugging Face06yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes6k downloads7mo agoHugging Face07cyjin-yl /m4singerForDiffSingerSVSaudio10K<n<100K1 likes5.4k downloads1y agoHugging Face08yasalma /tat_youtubeaudiotext-to-speech100K<n<1M0 likes5.3k downloads1y agoHugging Face09TheAgenticDataCompany /open-yap-1k Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0, listenable in the dataset viewer. The full corpus - 1,000 hours, 1,602… See the full description on the dataset page: https://huggingface.co/datasets/TheAgenticDataCompany/open-yap-1k.audioaudio-to-audion<1K105 likes5.1k downloads19d agoHugging Face10Vyvo-Research /Emilia-YODAS-ENaudio10M<n<100M4 likes3.8k downloads11mo agoHugging Face11ming030890 /youtube_caption_yue YouTube ASR Caption Dataset (Cantonese) This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re-transcribe the audio and filtered segments to build a high-quality collection of audio-caption pairs. What’s included Segments where the ASR output is identical to the original caption — likely clean. Segments where differences are only homophones (同音字) or English words — likely ASR mistakes. This combination supports… See the full description on the dataset page: https://huggingface.co/datasets/ming030890/youtube_caption_yue.audio10K<n<100K2 likes3.5k downloads1y agoHugging Face12TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.1k downloads2y agoHugging Face132e8konjak /youtube_audios_2audio0 likes2.8k downloads1y agoHugging Face14overflowwwww /yt-danish-public-v2audioaudio-classification100K<n<1M0 likes2.7k downloads2y agoHugging Face15retkowski /ytseg YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.audiotoken-classification100K<n<1M8 likes2.7k downloads2mo agoHugging Face16yinanli1215 /Sound2Hap[Last Update Feb 6 2026] The 1000 sound clips (audio1000 folder) are from:ESC-50: Dataset for Environmental Sound Classificationhttps://github.com/karoldvl/ESC-50/@inproceedings{piczak2015dataset, title = {{ESC}: {Dataset} for {Environmental Sound Classification}}, author = {Piczak, Karol J.}, booktitle = {Proceedings of the 23rd {Annual ACM Conference} on {Multimedia}}, date = {2015-10-13}, url = {http://dl.acm.org/citation.cfm?doid=2733373.2806390}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/yinanli1215/Sound2Hap.audio1K<n<10K0 likes2k downloads27d agoHugging Face17ylacombe /english_dialects Dataset Card for "english_dialects" Dataset Summary This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English. The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.audiotext-to-speech10K<n<100K37 likes1.9k downloads3y agoHugging Face18YirongSun /SonicBench SonicBench: Dissecting the Physical Perception Bottleneck in Large Audio Language Models 12 physical attributes, 5 perceptual dimensions, 2 task types - dissecting the physical perception bottleneck of Large Audio Language Models. Benchmark • Directory Layout • JSON Format • Probe Splits • Paper TL;DR. SonicBench is a psychophysically grounded benchmark that probes physical audio perception rather than semantics:12… See the full description on the dataset page: https://huggingface.co/datasets/YirongSun/SonicBench.audio1K<n<10K0 likes1.7k downloads8mo agoHugging Face19yhytoto12 /behavior-sd 🎙️ Behavior-SD Official repository for our NAACL 2025 paper:Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language ModelsSehun Lee*, Kang-wook Kim*, Gunhee Kim (* Equal contribution) 🏆 SAC Award Winner in Speech Processing and Spoken Language Understanding 🔗 Links Project Page Code 📖 Overview We explores how to generate natural, behaviorally-rich full-duplex spoken dialogues using large language models (LLMs). We introduce:… See the full description on the dataset page: https://huggingface.co/datasets/yhytoto12/behavior-sd.audio100K<n<1M9 likes1.7k downloads1y agoHugging Face202e8konjak /youtube_audios_11audio0 likes1.7k downloads1y agoHugging Face21ylacombe /expresso The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided. You can… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/expresso.audio10K<n<100K99 likes1.7k downloads2y agoHugging Face22YomnaGharib /dahih-tts2-demucs-cleanedaudio10K<n<100K1 likes1.6k downloads4mo agoHugging Face23YuxiangW /VoxPrivacy VoxPrivacy This repository contains a Parquet-packaged version of VoxPrivacy. Configs train contains the conversation examples. All source files are part of this config and are distinguished by split: zh_nobody_three_round zh_onlyme_three_round zh_no_secret en_nobody_three_round en_onlyme_three_round en_no_secret zh_no_secret2 Each example row has: id: stable row id split: split name task_type: original task type example: compact JSON string of the original row… See the full description on the dataset page: https://huggingface.co/datasets/YuxiangW/VoxPrivacy.audio1M<n<10M0 likes1.5k downloads4mo agoHugging Face24yihao005 /Multi-Talker-SD Dataset Card for Multi-Talker-SD Dataset Description Multi-Talker-SD is a large-scale bilingual (English–Mandarin) multi-speaker meeting dataset designed to support research on speaker diarization and meeting transcription. Size: 1,000 simulated meetings Participants per meeting: 10–30 speakers Average duration: ~20 minutes per meeting, up to one hour Languages: English, Mandarin (code-switching possible) Audio characteristics: realistic speaker overlap… See the full description on the dataset page: https://huggingface.co/datasets/yihao005/Multi-Talker-SD.audioautomatic-speech-recognition4 likes1.4k downloads1y agoHugging Face25Yusen /musicaudio1K<n<10K0 likes1.3k downloads1mo agoHugging Face26jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.3k downloads4y agoHugging Face27NCSpeech /YO-CPT-kk YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.audiotext-to-speech100K<n<1M10 likes1.3k downloads2mo agoHugging Face28naloz /YMlibertyaudion<1K0 likes1.3k downloads8d agoHugging Face29mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.3k downloads3y agoHugging Face30Yousra0x0 /LUNA16_RAWaudion<1K0 likes1.2k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.