CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ziggylott /tlott-digital-products T. Lott Digital Products Digital product files for T. Lott's online store. Products Audiobooks (MP3) eBooks (PDF) Software (ZIP) Cover images (PNG) Download URLs Files can be downloaded directly: https://huggingface.co/datasets/ziggylott/tlott-digital-products/resolve/main/{filepath} audion<1K0 likes5.2k downloads21d agoHugging Face02anke01 /uyghur-common-voice-tts Uyghur Common Voice TTS Dataset A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice. Dataset Summary Property Value Language Uyghur (ug) Total Samples 43,054 Train Samples 40,901 Validation Samples 2,153 Audio Format WAV Source Mozilla Common Voice License CC0-1.0 Dataset Structure / ├── train.jsonl # Training data (40,901 samples) ├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.audiotext-to-speech10K<n<100K0 likes3.9k downloads7mo agoHugging Face03thanhnew2001 /VietSuperSpeech VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.audio10K<n<100K6 likes3.3k downloads7mo agoHugging Face04FreedomIntelligence /TalkVid TalkVid Dataset This repository hosts the TalkVid dataset. Paper: TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis Arxiv paper: https://arxiv.org/abs/2508.13618 Project Page: https://freedomintelligence.github.io/talk-vid GitHub: https://github.com/FreedomIntelligence/TalkVid Abstract Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TalkVid.audioimage-to-videon<1K22 likes3k downloads1y agoHugging Face05tutu0604 /UltraVoice UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset… See the full description on the dataset page: https://huggingface.co/datasets/tutu0604/UltraVoice.audiotext-to-speech100K<n<1M16 likes2.2k downloads11mo agoHugging Face06YomnaGharib /dahih-tts2-demucs-cleanedaudio10K<n<100K1 likes1.4k downloads4mo agoHugging Face07titasmallick96 /daily-bio-newsaudion<1K0 likes1.4k downloads2h agoHugging Face08jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.3k downloads4y agoHugging Face09wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes802 downloads1mo agoHugging Face10McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes762 downloads1mo agoHugging Face11Harland /DCASE2026-Task5-DevSet DCASE 2026 Task 5 Audio-Dependent Question Answering (ADQA) Development Set This is the official Development Set for DCASE 2026 Challenge Task 5: Audio-Dependent Question Answering (ADQA). The ADQA task focuses on addressing "Textual Hallucination" in Large Audio-Language Models (LALMs) — where models pass audio understanding benchmarks by relying on text prompts and internal linguistic priors rather than actual audio perception. ADQA introduces a rigorous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Harland/DCASE2026-Task5-DevSet.audio1K<n<10K1 likes564 downloads2mo agoHugging Face12leungtianle /AgentChat-Test Test Set Description This directory contains the test set used for tool-use evaluation. The JSON files under Test-JSON/ are organized by task type: SingleTaskProcessing/tool-select_test.json: single-tool selection tasks. ParallelProcessing/parallel-call_test.json: parallel tool-call tasks. ProactiveSeeking/searchTools_test_predictions_kept.json: proactive tool-search tasks. TaskDecomposition/muti-tool-select_test.json: multi-tool task decomposition tasks.… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/AgentChat-Test.audion<1K0 likes393 downloads3mo agoHugging Face13yuanzhuyun /asr-reference-set-eval-temp Temporary ASR evaluation audio Temporary public audio files used for hosted ASR evaluation. audio1K<n<10K0 likes357 downloads2mo agoHugging Face14nymtheescobar /bengali-talkshow-audio Bengali Talkshow Audio Dataset A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs. Dataset Description This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.audioaudio-classification1K<n<10K0 likes294 downloads8mo agoHugging Face15ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes287 downloads5mo agoHugging Face16whyismydininghallonfire /orig-plus-asr-tamil-clean orig-plus-asr-tamil-clean Combined ASR dataset built from: albagon/til26-asr-split (orig rows) whyismydininghallonfire/asr-tamil-clean (asr_tamil_clean rows) Audio paths are namespaced under each split to avoid filename collisions: audio/orig/... audio/asr_tamil_clean/... Each row keeps key, audio, transcript, and language, with an added source_dataset field. Counts: train: 3595 orig + 891 asr_tamil_clean = 4486 validation: 899 orig + 224 asr_tamil_clean = 1123 audio1K<n<10K0 likes216 downloads4mo agoHugging Face17mcp-tool-shop /jam-actions-v1 jam-actions-v1 Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) · Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE · Source repo: mcp-tool-shop-org/ai-jam-sessions The successor to jam-actions-v0. Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from what the tools return — and it exists in its current shape because, seven training runs in a row, the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.texttext-generationn<1K0 likes206 downloads16d agoHugging Face18taohu /music2chords_v2audio1K<n<10K0 likes190 downloads7mo agoHugging Face19panlr /teochew_wildgated Teochew-Wild:首个正字标注的野外潮州话数据集 本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正); Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。 文件说明 (File Structure Explanation) ├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式 ├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.audiotext-to-speech10K<n<100K45 likes179 downloads2mo agoHugging Face20ICTNLP /MutiEmo-Test MultiEmo-Test MultiEmo-Test is an English evaluation set for instruction-following multi-emotion text-to-speech synthesis. It accompanies HybridEmo, a system for modeling sequential emotion trajectories and simultaneous emotion blending within an utterance. The dataset is intended for evaluation only. It contains synthesis text, natural-language emotion instructions, emotion annotations, and prompt audio for speaker-timbre conditioning. It does not contain target synthesized… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/MutiEmo-Test.audion<1K1 likes175 downloads24d agoHugging Face21zsy814 /instructtts-three-model-gemini-zh InstructTTSEval 三模型 Gemini 评测数据 本目录整理了 InstructTTSEval 中文集上三个 TTS 模型的生成音频和 Gemini 一致性评测结果:Qwen3-TTS-12Hz-1.7B-VoiceDesign、Seed-Audio-1.0、VoxCPM2。 字段 records.jsonl 每行对应一个模型和一种控制格式(APS、DSD 或 RP): id:InstructTTSEval 样本 ID mode:控制格式 model、model_name:模型标识 text:合成文本 instruction:历史评测记录中的输入控制指令,按本行 APS/DSD/RP 格式保留;不等同于各模型 API 的完整请求封装 generated_audio:该模型生成音频的相对路径 reference_audio:原始参考音频的相对路径 gemini_consistent:Gemini judge 的一致性判断 inconsistency_reason:判断为不一致时的原因… See the full description on the dataset page: https://huggingface.co/datasets/zsy814/instructtts-three-model-gemini-zh.audio1K<n<10K0 likes163 downloads9d agoHugging Face22mcp-tool-shop /jam-actions-v1-probe jam-actions-v1-probe Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only · Companion to: jam-actions-v1 Why it exists An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes. Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.textothern<1K0 likes160 downloads14d agoHugging Face23mcp-tool-shop /jam-actions-acoustic-v0 Dataset Card for jam-actions-acoustic-v0 Version: 1.0.2 Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI. Summary 108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence). This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.texttext-generationn<1K0 likes158 downloads17d agoHugging Face24danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes145 downloads10mo agoHugging Face25Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes143 downloads1y agoHugging Face26tterumiimurett1 /agentic-asrgated Agentic ASR Public consolidated audio and ASR result dataset for the OSWorld and WildClawBench benchmark families. Layout osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise pairs, task images, ASR results, and reports. wildclawbench/: 60 formal colloquialized prompts, synthetic speech, 20 synthetic ASR condition tables, and ten-participant human recordings. task0_template derivatives are excluded. metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.audio10K<n<100K0 likes140 downloads4d agoHugging Face27ai-ssam /darija-tts-8400 Darija TTS 8400 Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV. All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings. Write-up of how this data was used: Training a Voice. At a glance Clips / hours 8,400 / 20.73 Unique texts 4,800 Voice Kore (1 speaker) Sample rate 24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.audiotext-to-speech1K<n<10K0 likes122 downloads9d agoHugging Face28teamsleeping /Kyrg-TTSaudio1K<n<10K0 likes100 downloads1mo agoHugging Face29tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes84 downloads1mo agoHugging Face30TNSA /Aren ARen — Arabic/English ASR Robustness Set Curated and published by TNSA AI. A small, deliberately hard evaluation set for Arabic and English speech recognition. Every clip exists in three acoustic conditions so you can measure not just how a model scores, but how fast it falls apart as the channel degrades. Built because clean read-speech benchmarks stop discriminating between modern ASR systems long before real deployments stop breaking. Why it exists On clean… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/Aren.audioautomatic-speech-recognitionn<1K0 likes81 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.