CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes1.6k downloads2mo agoHugging Face02martinturuta /safi-kinyarwanda-conversations Safi Diction Kinyarwanda Conversational Speech Dataset This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine. The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.audioautomatic-speech-recognitionn<1K0 likes336 downloads19d agoHugging Face03HTH-inc /japanese-casual-conversational-speech-golden-dataset-preview Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.audioautomatic-speech-recognitionn<1K2 likes238 downloads22d agoHugging Face04Makan09 /bam-asr-conversational All Bambara ASR Dataset This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-conversational.audioautomatic-speech-recognition10K<n<100K2 likes234 downloads24d agoHugging Face05FormosanBank /ePark_sheng_huo_hui_hua_pian_daily_conversation FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation.audioautomatic-speech-recognition10K<n<100K0 likes217 downloads2mo agoHugging Face06AirCaps /mega-asr-conversational-overlap Mega-ASR Conversational Overlap Mega-ASR Conversational Overlap is a deterministic English ASR diagnostic set derived from AirCaps/mega-asr-noise-a5sv2, which in turn is sampled from the Mega-ASR training corpus zhifeixie/Voices-in-the-Wild-2M. The existing AirCaps dataset evaluates single-utterance acoustic robustness. This companion dataset evaluates a different failure mode: two-turn conversational continuity with slight overlap and unequal turn loudness. It does not replace… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-conversational-overlap.audioautomatic-speech-recognitionn<1K0 likes150 downloads29d agoHugging Face07jml2026 /conversational-speech-dataset 🎙️ Silencio Network: Conversational Speech Dataset Overview Sample conversational speech data from Silencio Network's crowdsourced voice AI platform. This dataset contains multi-speaker meeting recordings with word-level transcripts, speaker diarization, and rich demographic metadata. Each row represents one participant in a meeting and includes 3 audio files: Audio Column Description Format file_name (speaker audio) Individual participant's… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/conversational-speech-dataset.automatic-speech-recognitionn<1K0 likes141 downloads6mo agoHugging Face08Appenlimited /1000h-us-english-smartphone-conversation 📚 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings) This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for: Automatic Speech Recognition (ASR) Speaker Identification and Gender/Age Analysis Dialect and Accent Modeling Multi-speaker Speech Separation 🧾 Dataset Contents The dataset includes: metadata.CSV: Metadata including speaker gender, age… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.audioautomatic-speech-recognitionn<1K3 likes138 downloads1y agoHugging Face09bbdontcry /en-everyday-conversation-asr English Everyday-Conversation ASR dataset 34777 clips · 39.996 hours · 16 kHz mono · CC-BY-SA 4.0 Assembled from two commercial-safe (CC-BY-SA 4.0) open corpora: EdAcc (spontaneous dyadic conversation, 11485 clips) and DailyTalk (scripted everyday-life dialogue, 23292 clips). Splits ({'train': 31297, 'dev': 3480}) are conversation-disjoint (no speaker leakage). Per-row metadata (gender, accent, style, l1, ...) is preserved so the set can be re-balanced. from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/bbdontcry/en-everyday-conversation-asr.audioautomatic-speech-recognition10K<n<100K0 likes137 downloads4mo agoHugging Face10BoxlyX /English_Natural_Conversation_ASR_STT BoxlyX English Natural Conversation Sample Dataset (ASR/STT) 📌 Overview This repository contains high-fidelity, studio-recorded English natural conversation samples designed for training and benchmarking advanced Automatic Speech Recognition (ASR) and Speech-to-Text (STT) models. This dataset is a curated public sample provided by BoxlyX AI Solution, showcasing our end-to-end capabilities in premium audio data generation, multi-speaker recording environment… See the full description on the dataset page: https://huggingface.co/datasets/BoxlyX/English_Natural_Conversation_ASR_STT.audioautomatic-speech-recognitionn<1K4 likes135 downloads3mo agoHugging Face11arcada-labs /conversation-bench Conversation Bench 75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.audioautomatic-speech-recognitionn<1K8 likes88 downloads6mo agoHugging Face12UniDataPro /human-robot-conversation-russian Human-Robot Dataset The dataset comprises 660+ hours of Russian speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-russian.audioautomatic-speech-recognitionn<1K1 likes84 downloads1mo agoHugging Face13beatsprom /realtime-conversational-voice-agent-duplex-2026 🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026) This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.texttext-generationn<1K0 likes69 downloads21d agoHugging Face14UniDataPro /human-robot-conversation-korean Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes53 downloads1mo agoHugging Face15Indospeech /sundanese-spontaneous-conversation-samplegated Sundanese Spontaneous Conversation — Free Sample Unscripted two-speaker Sundanese (su-ID) conversation recorded in West Java, Indonesia. This is a free 30-minute sample of a larger commercially licensable corpus. Why this exists Sundanese has approximately 40 million native speakers, yet spontaneous conversational data is almost absent: Resource Type Volume Common Voice Spontaneous 4.0 Spontaneous, 2 speakers 0.62 h OpenSLR SLR36 / SLR44 Read speech —… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.audioautomatic-speech-recognitionn<1K1 likes52 downloads10d agoHugging Face16ud-nlp /human-robot-conversation-korean Human-Robot Conversation Dataset (Korean) - 660+ Hours Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes48 downloads6mo agoHugging Face17UniDataPro /human-robot-conversation-german Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.audioautomatic-speech-recognitionn<1K1 likes46 downloads1mo agoHugging Face18ud-nlp /human-robot-conversation-russian Human-Robot Conversation Dataset (Russian) - 660+ Hours Dataset (Russian) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-russian.audioautomatic-speech-recognitionn<1K1 likes44 downloads6mo agoHugging Face19snorbyte /indic-audio-natural-conversations-samplegated Dataset Card for Indic Audio Natural Conversations Sample Dataset Dataset Details Dataset Description The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.audioaudio-to-audion<1K1 likes38 downloads1y agoHugging Face20Ethan615 /taiwan-conversation-context-100-domainsgated Taiwan Conversation Context 100 Domains Dataset Description Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。 本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。 資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於: 語音生成資料前處理 Text-to-Speech, TTS Spoken Dialogue Generation Conversational AI Customer Service Dialogue Modeling Role-play Dialogue Dataset 台灣繁體中文語音模型訓練 生活情境問答模型訓練 對話式 AI 助理訓練 RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.texttext-generation1M<n<10M2 likes38 downloads5mo agoHugging Face21Zeldeo /french-conversation_structured Zeldeo/french-conversation_structured Dataset ASR restructuré depuis Snit/french-conversation (config=default, split=train). Nombre d'exemples : 98. Métadonnées ajoutées : source_dataset, type, langue_accent. Normalisation texte : aucune. Colonnes conservées audio transcription id part audio_path Usage from datasets import load_dataset ds = load_dataset("Zeldeo/french-conversation_structured", split="train") print(ds[0]) audioautomatic-speech-recognitionn<1K0 likes38 downloads29d agoHugging Face22deepdml /conversations ATC Pilot–Controller Conversations Dataset Description This dataset contains English conversations between pilots and air traffic controllers (ATC), transcribed from real VHF radio transmissions recorded at controlled aerodromes. Each example consists of: A WAV audio file (mono, VHF radio quality) A transcription of the utterance Speaker metadata: whether the speaker is the pilot or the controller Utterance type: e.g. readback, taxi, pushback, etc. Token-level… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/conversations.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads7mo agoHugging Face23UniDataPro /human-robot-conversation-english Human-Robot Dataset The dataset comprises 660+ hours of English speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-english.audioautomatic-speech-recognitionn<1K0 likes35 downloads1mo agoHugging Face24Datoric /tts-conversational-voice-20000hgated TTS Voice Dataset 20,000 hours of high-fidelity 48kHz conversational audio across 30+ global, regional, and underrepresented languages, built for text-to-speech, voice cloning, and multilingual speech AI. This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real audio samples. Overview… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/tts-conversational-voice-20000h.tabulartext-to-speech100K<n<1M0 likes35 downloads3mo agoHugging Face25ud-nlp /human-robot-conversation-english Human-Robot Conversation Dataset (English) - 660+ Hours Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.audioautomatic-speech-recognitionn<1K1 likes28 downloads6mo agoHugging Face26ShimogaAIteam /conversational_kannada_stt Conversational Kannada STT This dataset contains corrected transcriptions of conversational Kannada speech,prepared specifically for fine-tuning Whisper models on conversational and dialectal Kannada. Unlike many ASR datasets, this release provides pre-computed Whisper input features (log-Mel spectrograms)so you can train/fine-tune Whisper models without raw audio processing. Dataset Creation Source Audio: Publicly available YouTube videos in Kannada. Initial… See the full description on the dataset page: https://huggingface.co/datasets/ShimogaAIteam/conversational_kannada_stt.textautomatic-speech-recognition1K<n<10K1 likes27 downloads1y agoHugging Face27Datoric /audio-video-conversation-4000hgated Audio-Video Conversational Dataset 4,000 hours of synchronized speech and video of natural conversations: face movement, mouth motion, gestures, turn-taking, emotion, laughter, and interruptions across 20+ languages. This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real samples.… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/audio-video-conversation-4000h.tabularautomatic-speech-recognition100K<n<1M0 likes25 downloads3mo agoHugging Face28UsergyAI /Global-Conversational-Speechgated Global Conversational Speech Dataset 305 hours. 18 locales. Real conversations. Not scraped from YouTube. Not recorded by anonymous crowds who don't speak the language. Every conversation in this dataset traces back to verified native speakers we know by name. The [Human] Standard Most speech datasets are built the same way: scrape the internet, hire anonymous contractors, run it through automated QC, ship it. The result? Models that are confidently wrong. We… See the full description on the dataset page: https://huggingface.co/datasets/UsergyAI/Global-Conversational-Speech.audioautomatic-speech-recognition100K<n<1M2 likes24 downloads2mo agoHugging Face29RyeAI /coral-v3-conversation-pnc-dagated CoRal v3 Conversation PnC DA RyeAI/coral-v3-conversation-pnc-da is a text-only Danish punctuation and capitalization companion for the conversation training split of CoRal-project/coral-v3. It contains 102,226 restored transcript rows and no audio bytes. Pinned companion revision: 6e4fbafde87fbffadd58bbe39a3a2e09e884351a. SHA-256 of data/train-00000-of-00001.parquet: 79d30671f238e884a5b71b682bc811043e7df075566ce5565a236e66ffe32000. Each row can be joined back to the gated source… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/coral-v3-conversation-pnc-da.tabularautomatic-speech-recognition100K<n<1M0 likes22 downloads18d agoHugging Face30DatoricAI /audio-video-conversation-4000hgated Audio-Video Conversational Dataset 4,000 hours of synchronized speech and video of natural conversations: face movement, mouth motion, gestures, turn-taking, emotion, laughter, and interruptions across 20+ languages. This repository is a specification and preview listing. The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real samples. Overview The Audio-Video Conversational Dataset is a 4… See the full description on the dataset page: https://huggingface.co/datasets/DatoricAI/audio-video-conversation-4000h.textautomatic-speech-recognitionn<1K0 likes18 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.