CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes1.6k downloads2mo agoHugging Face02JDKdev /french-tts-conversational-dataset French Conversational TTS Dataset Dataset Description This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution. Verticals Vertical Description fintech_banking Banking operations, account inquiries, fraud alerts, investments, customer service ecommerce_logistics Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.audiotext-to-speechn<1K0 likes910 downloads3mo agoHugging Face03nytopop /expresso-conversational The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided. You can listen to… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/expresso-conversational.audio10K<n<100K14 likes424 downloads1y agoHugging Face04MagicHub /multi-stream-spontaneous-conversation-training-datasets_chinese Multi-stream Spontaneous Conversation Training Datasets_Chinese Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: Chinese Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_chinese.audio1K<n<10K2 likes372 downloads3mo agoHugging Face05martinturuta /safi-kinyarwanda-conversations Safi Diction Kinyarwanda Conversational Speech Dataset This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine. The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.audioautomatic-speech-recognitionn<1K0 likes339 downloads19d agoHugging Face06khamidov17 /clinical-conversations-anon-benchmarkaudion<1K0 likes324 downloads3mo agoHugging Face07MagicDataTech /multi-stream-spontaneous-conversation-training-datasets_chinese Multi-stream Spontaneous Conversation Training Datasets_Chinese Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: Chinese Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/multi-stream-spontaneous-conversation-training-datasets_chinese.audio1K<n<10K2 likes265 downloads4mo agoHugging Face08HTH-inc /japanese-casual-conversational-speech-golden-dataset-preview Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.audioautomatic-speech-recognitionn<1K2 likes246 downloads22d agoHugging Face09Makan09 /bam-asr-conversational All Bambara ASR Dataset This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-conversational.audioautomatic-speech-recognition10K<n<100K2 likes236 downloads24d agoHugging Face10MagicHub /korean-conversational-speech-corpus ASR-KCSC: A Korean Conversational Speech Corpus Every data point counts. Dataset Basic Info Dataset Type: ASR Speech Corpus Language: Korean Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Recording Environment: Indoor Dataset Description This open-source dataset consists of 5.22 hours of transcribed Korean conversational speech on certain topics, where 22 conversations between seven pairs of speakers… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/korean-conversational-speech-corpus.audio1K<n<10K1 likes223 downloads3mo agoHugging Face11FormosanBank /ePark_sheng_huo_hui_hua_pian_daily_conversation FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation.audioautomatic-speech-recognition10K<n<100K0 likes208 downloads2mo agoHugging Face12niloy629 /personaplex-distill-conversations PersonaPlex Distillation Conversation Dataset Teacher-generated multi-turn conversation data for distilling/pruning NVIDIA PersonaPlex 7B (a Moshi-architecture full-duplex speech-to-speech model). What this is Each sample is a real conversation rendered by the PersonaPlex teacher itself: The student's turns are scripted (generated by Qwen3-8B) and voiced with Piper TTS The teacher's (PersonaPlex's) responses are improvised live by the model — its real… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/personaplex-distill-conversations.audioaudio-to-audion<1K0 likes179 downloads12d agoHugging Face13MagicHub /multi-stream-spontaneous-conversation-training-datasets_english Multi-stream Spontaneous Conversation Training Datasets_English Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: English Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_english.audio1K<n<10K1 likes175 downloads3mo agoHugging Face14lu-vae /Miracle-ConversationThis is the dataset curated from ChatGPT with personalized prompt from Our EMNLP2023-findings Miracle We offer three personality aspects: 'a' = attitude (positive/negative) 'l' = language style (lyrical/plain) 'm' = mental characteristics (critical/emotional) audio1 likes161 downloads2y agoHugging Face15TingChen-ppmc /Shanghai_Dialect_Conversational_Speech_Corpus Corpus This dataset is built from Magicdata ASR-CZDIACSC: A CHINESE SHANGHAI DIALECT CONVERSATIONAL SPEECH CORPUS This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information. Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed. Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Shanghai_Dialect_Conversational_Speech_Corpus.audio1K<n<10K12 likes160 downloads2y agoHugging Face16Snit /french-conversation+15 hours of speech data from TTS and text file recording. +9k utterances from various sources, novels, parliamentary debates, professional language. audio10K<n<100K6 likes146 downloads3y agoHugging Face17Appenlimited /1000h-us-english-smartphone-conversation 📚 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings) This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for: Automatic Speech Recognition (ASR) Speaker Identification and Gender/Age Analysis Dialect and Accent Modeling Multi-speaker Speech Separation 🧾 Dataset Contents The dataset includes: metadata.CSV: Metadata including speaker gender, age… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.audioautomatic-speech-recognitionn<1K3 likes137 downloads1y agoHugging Face18bbdontcry /en-everyday-conversation-asr English Everyday-Conversation ASR dataset 34777 clips · 39.996 hours · 16 kHz mono · CC-BY-SA 4.0 Assembled from two commercial-safe (CC-BY-SA 4.0) open corpora: EdAcc (spontaneous dyadic conversation, 11485 clips) and DailyTalk (scripted everyday-life dialogue, 23292 clips). Splits ({'train': 31297, 'dev': 3480}) are conversation-disjoint (no speaker leakage). Per-row metadata (gender, accent, style, l1, ...) is preserved so the set can be re-balanced. from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/bbdontcry/en-everyday-conversation-asr.audioautomatic-speech-recognition10K<n<100K0 likes137 downloads4mo agoHugging Face19alakxender /dhivehi-conversations-turn Dhivehi Conversations (Turn-Based) This is an experimental synthetic dataset of turn-based Dhivehi conversations created for testing and fine-tuning dialogue models, text-to-speech (TTS), and multi-turn speaker-aware systems. This dataset is artificially constructed and not based on real conversations. It is intended for research experimentation only and may not always produce contextually accurate results. Dataset Source Derived from alakxender/voice-synthetic… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-conversations-turn.audiotext-to-speech10K<n<100K0 likes136 downloads1y agoHugging Face20BoxlyX /English_Natural_Conversation_ASR_STT BoxlyX English Natural Conversation Sample Dataset (ASR/STT) 📌 Overview This repository contains high-fidelity, studio-recorded English natural conversation samples designed for training and benchmarking advanced Automatic Speech Recognition (ASR) and Speech-to-Text (STT) models. This dataset is a curated public sample provided by BoxlyX AI Solution, showcasing our end-to-end capabilities in premium audio data generation, multi-speaker recording environment… See the full description on the dataset page: https://huggingface.co/datasets/BoxlyX/English_Natural_Conversation_ASR_STT.audioautomatic-speech-recognitionn<1K4 likes136 downloads3mo agoHugging Face21AirCaps /mega-asr-conversational-overlap Mega-ASR Conversational Overlap Mega-ASR Conversational Overlap is a deterministic English ASR diagnostic set derived from AirCaps/mega-asr-noise-a5sv2, which in turn is sampled from the Mega-ASR training corpus zhifeixie/Voices-in-the-Wild-2M. The existing AirCaps dataset evaluates single-utterance acoustic robustness. This companion dataset evaluates a different failure mode: two-turn conversational continuity with slight overlap and unequal turn loudness. It does not replace… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-conversational-overlap.audioautomatic-speech-recognitionn<1K0 likes113 downloads1mo agoHugging Face22malaysia-ai /malay-conversational-speech-corpus malay-conversational-speech-corpus Mirror for https://magichub.com/datasets/malay-conversational-speech-corpus/, license is Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License audio1K<n<10K5 likes110 downloads3y agoHugging Face23TingChen-ppmc /Changsha_Dialect_Conversational_Speech_Corpus Corpus This dataset is built from Magicdata ASR-CCHSHDIACSC: A CHINESE CHANGSHA DIALECT CONVERSATIONAL SPEECH CORPUS This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information. Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed.… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Changsha_Dialect_Conversational_Speech_Corpus.audio1K<n<10K2 likes103 downloads3y agoHugging Face24TingChen-ppmc /Nanchang_Dialect_Conversational_Speech_Corpus Corpus This dataset is built from Magicdata ASR-CNANDIACSC: A CHINESE NANCHANG DIALECT CONVERSATIONAL SPEECH CORPUS This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information. Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed. Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Nanchang_Dialect_Conversational_Speech_Corpus.audio1K<n<10K1 likes96 downloads3y agoHugging Face25eustlb /dailytalk-conversations-grouped Dataset Card for "dailytalk-conversations-grouped" This dataset is intended for testing fine-tuning of Sesame’s CSM-1B (Conversational Speech Model). It is extracted from the DailyTalk dataset. audio1K<n<10K11 likes96 downloads1y agoHugging Face26arcada-labs /conversation-bench Conversation Bench 75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.audioautomatic-speech-recognitionn<1K8 likes92 downloads6mo agoHugging Face27equal-ai /conversational_englishaudio10K<n<100K0 likes89 downloads7mo agoHugging Face28UniDataPro /human-robot-conversation-russian Human-Robot Dataset The dataset comprises 660+ hours of Russian speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-russian.audioautomatic-speech-recognitionn<1K1 likes87 downloads1mo agoHugging Face29Rcarvalo /conversational-s2s-v1 Conversational S2S v1 — dialogues parlés FR + EN Corpus synthétique de conversations orales multi-tours entre un utilisateur et un assistant, en français et en anglais, destiné au finetuning speech-to-speech de LFM2.5-Audio (Liquid AI). Chaque tour est un clip audio séparé, aligné avec son texte ; les dialogues sont conçus pour être rejoués tour à tour (user → assistant → user → …). Projet tts-model-exploration. Produit par voxtral_datagen_pipeline (s2s-skeletons → remplissage… See the full description on the dataset page: https://huggingface.co/datasets/Rcarvalo/conversational-s2s-v1.audioaudio-to-audio10K<n<100K0 likes79 downloads13d agoHugging Face30TingChen-ppmc /Zhengzhou_Dialect_Conversational_Speech_Corpus Corpus This dataset is built from Magicdata ASR-CZDIACSC: A CHINESE ZHENGZHOU DIALECT CONVERSATIONAL SPEECH CORPUS This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information. Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed. Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Zhengzhou_Dialect_Conversational_Speech_Corpus.audio1K<n<10K3 likes75 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.