CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B89 likes8.2k downloads6mo agoHugging Face02tutu0604 /UltraVoice UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset… See the full description on the dataset page: https://huggingface.co/datasets/tutu0604/UltraVoice.audiotext-to-speech100K<n<1M16 likes2.2k downloads11mo agoHugging Face03wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes802 downloads1mo agoHugging Face04McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes762 downloads1mo agoHugging Face05Quran-Lab /quran-tajweed-phonetics The complete phonetic layer of the Quran in the riwaya of Hafs 'an 'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every phone carrying its tajweed attribution: madd class with its transmitted length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt, the seventeen sifat, and the rule that produced it. Built and maintained by Quran Lab, a waqf building open technology in the service of the Quran. How it was built and verified Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.tabularautomatic-speech-recognition10K<n<100K3 likes532 downloads13d agoHugging Face06pykeio /librivox-tracksA dataset of all audio files & corresponding sources uploaded to LibriVox before 1st November 2025. Donate to Internet Archive, who hosts most of the audio data Donate to Project Gutenberg, who hosts most of the source books Be a good netizen and use proper caching & sensible rate limiting when downloading files. November 2025 Update book.id, section.id, and section.num are now integers instead of strings. reader has been replaced with a readers array. No entries currently… See the full description on the dataset page: https://huggingface.co/datasets/pykeio/librivox-tracks.texttext-to-speech100K<n<1M13 likes327 downloads11mo agoHugging Face07nymtheescobar /bengali-talkshow-audio Bengali Talkshow Audio Dataset A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs. Dataset Description This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.audioaudio-classification1K<n<10K0 likes294 downloads8mo agoHugging Face08ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes287 downloads5mo agoHugging Face09panlr /teochew_wildgated Teochew-Wild:首个正字标注的野外潮州话数据集 本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正); Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。 文件说明 (File Structure Explanation) ├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式 ├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.audiotext-to-speech10K<n<100K45 likes179 downloads2mo agoHugging Face10thepowerfuldeez /massive-yt-edu-queue Massive YouTube Educational Video Queue Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours. Description This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.tabularautomatic-speech-recognition1M<n<10M1 likes146 downloads7mo agoHugging Face11danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes145 downloads10mo agoHugging Face12Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes143 downloads1y agoHugging Face13oncody /AI_Agent_Task_Dataset 🤖 Massive AI Agent Task Dataset (10.5GB) 📌 Overview Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs. This dataset focuses on: Multi-step reasoning Tool usage (APIs, frameworks, systems) Real-world execution workflows Perfect for building agentic AI systems, copilots, and automation models. 📑 Table of Contents Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.texttext-generation10M<n<100M3 likes85 downloads6mo agoHugging Face14tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes84 downloads1mo agoHugging Face15TNSA /Aren ARen — Arabic/English ASR Robustness Set Curated and published by TNSA AI. A small, deliberately hard evaluation set for Arabic and English speech recognition. Every clip exists in three acoustic conditions so you can measure not just how a model scores, but how fast it falls apart as the channel degrades. Built because clean read-speech benchmarks stop discriminating between modern ASR systems long before real deployments stop breaking. Why it exists On clean… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/Aren.audioautomatic-speech-recognitionn<1K0 likes81 downloads1mo agoHugging Face16mesolitica /pseudolabel-malaya-speech-stt-train-whisper-large-v3tabularautomatic-speech-recognition1M<n<10M1 likes61 downloads3y agoHugging Face17rustam1221 /uzbek-asr-train-manifests Uzbek ASR Training Manifests The exact training, validation and test splits behind rustam1221/uzbek-asr-gigaam: 974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized, and split by speaker. No audio is copied. Each row is a pointer — a parquet file plus a row index in the upstream dataset — and the training dataloader decodes the audio when the batch is built. That keeps the whole corpus definition at 200 MB instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.textautomatic-speech-recognition1K<n<10K0 likes59 downloads22d agoHugging Face18taras-sereda /uk-pods uk-pods - speech datasets of Ukrainian podcasts. Preparation Clone the dataset repository and extract the content of clips.tar.gz archive. git clone https://huggingface.co/datasets/taras-sereda/uk-pods cd uk-pods && tar -zxvf clips.tar.gz To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step. # data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods data_root=$(realpath .) sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.audioautomatic-speech-recognition10K<n<100K1 likes58 downloads2y agoHugging Face19TalTechNLP /err-video-news-transcribed Transcribed ERR Video News Dataset This dataset contains transcriptions of video news stories from Estonian National Brroacasting (https://www.err.ee/). There are around 40K stories with a total duration of around 4000 hours. Transcriptions are generated automatically using speech recognition (gemini-3-flash-preview). Contextual biasing was used to improve ASR quality, using the textual news story about the same topic. The WER of the transcriptions is around 5% on the average. The… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/err-video-news-transcribed.textautomatic-speech-recognition10K<n<100K1 likes50 downloads6mo agoHugging Face20danieldzikunuofmarvel /bibletts-asante-twi-repaired BibleTTS Asante Twi — Repaired Transcripts The Asante Twi transcripts released with BibleTTS have had the characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them. Audio is not included. This is a drop-in replacement for the .txt files that ship with the BibleTTS Asante Twi package, matched by clip ID. The problem Both are Twi vowels, and both are required by the orthography. Measured across the released Asante Twi transcripts: Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.tabularautomatic-speech-recognition10K<n<100K0 likes37 downloads2mo agoHugging Face21Codyfederer /test321 test321 This is a merged speech dataset containing 118 audio segments from 2 source datasets. Dataset Information Total Segments: 118 Speakers: 4 Languages: tr Emotions: happy, angry, sad, neutral Original Datasets: 2 Dataset Structure Each example contains: audio: Audio file (WAV format, 16kHz sampling rate) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged datasets) emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test321.audioautomatic-speech-recognitionn<1K0 likes31 downloads1y agoHugging Face22bdx33 /ted-talks-in-chinese-zhongwen TED中文 Podcast 聚焦华语地区的创意,本节目从上万个TED和TEDx演讲中,为您精选中文演讲,以及少量中文配音的经典英语演讲。演讲人包括科技和人文专家、关心当下与未来的思考者、关注挑战与探索的实践者。英雄不论出处,谁有创意谁讲。让这些演讲成为一把把钥匙,开启你的好奇心,升级你的行动力。 Focusing on creativity within the Chinese-speaking world, this program curates Chinese-language talks from tens of thousands of TED and TEDx presentations, along with a select few English classics dubbed into Chinese. Our speakers span technology and humanities experts, thinkers engaged with the present and future, and practitioners… See the full description on the dataset page: https://huggingface.co/datasets/bdx33/ted-talks-in-chinese-zhongwen.texttext-generationn<1K1 likes30 downloads1y agoHugging Face23theblackcat102 /quantized-common-voice-entextautomatic-speech-recognition1M<n<10M1 likes26 downloads3y agoHugging Face24theblackcat102 /common-voice-en-revoicetextautomatic-speech-recognition10K<n<100K0 likes22 downloads3y agoHugging Face25bingbangboom /cleaned-asr-transcriptstexttext-generation10K<n<100K1 likes22 downloads6mo agoHugging Face26zyzsasasas /testpic01 EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. Dataset Summary Languages: 22 European languages (see detailed breakdown below) Total aligned hours: ~78,100 hours of… See the full description on the dataset page: https://huggingface.co/datasets/zyzsasasas/testpic01.textautomatic-speech-recognitionn<1K0 likes18 downloads1y agoHugging Face27Boxp /ts_asr_test ts_asr_test:目标说话人 ASR 测试集(manifest-only) ts_asr_test is a 3,928-clip (~8.7 h) Chinese/English target-speaker ASR test set, released manifest-only: the repo ships no audio, only an audio-free recipe and a self-contained, deterministic rebuild script. Bring your own copies of the public source corpora and run rebuild_ts_asr_test.py to regenerate every clip bit-for-bit. 数据集简介 每条样本由一段目标说话人语音、一段同说话人的注册音频(enrollment),以及若干干扰说话人语音与背景噪声按固定配方混合而成。任务:在给定 enrollment… See the full description on the dataset page: https://huggingface.co/datasets/Boxp/ts_asr_test.textautomatic-speech-recognition1K<n<10K0 likes18 downloads4mo agoHugging Face28bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads5mo agoHugging Face29NbAiLab /freddy-testDette er et datasett som skal slettes. textautomatic-speech-recognition1K<n<10K0 likes15 downloads11mo agoHugging Face30SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes15 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.