datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.french-tts-conversational-dataset
French Conversational TTS Dataset
Dataset Description
This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution.
Verticals
Vertical
Description
fintech_banking
Banking operations, account inquiries, fraud alerts, investments, customer service
ecommerce_logistics
Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.expresso-conversational
The Expresso Dataset
[paper] [demo samples] [Original repository]
Introduction
The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided.
You can listen to… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/expresso-conversational.multi-stream-spontaneous-conversation-training-datasets_chinese
Multi-stream Spontaneous Conversation Training Datasets_Chinese
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Corpus
Language: Chinese
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Dataset Description
The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_chinese.safi-kinyarwanda-conversations
Safi Diction Kinyarwanda Conversational Speech Dataset
This dataset contains 1 hour of Kinyarwanda conversational speech collected using Safi's collection engine.
The recordings contain multiple speakers responding to survey questions. The original recordings were processed using speaker diarization to identify speaker turns. Consecutive turns from the same speaker were consolidated and split into speaker-specific audio clips of up to 15 seconds. These clips were then… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-kinyarwanda-conversations.clinical-conversations-anon-benchmarkmulti-stream-spontaneous-conversation-training-datasets_chinese
Multi-stream Spontaneous Conversation Training Datasets_Chinese
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Corpus
Language: Chinese
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Dataset Description
The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/multi-stream-spontaneous-conversation-training-datasets_chinese.japanese-casual-conversational-speech-golden-dataset-preview
Japanese Casual Conversational Speech Golden Dataset (Preview)
💼 Commercial License & Full Access
This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.
To purchase the full dataset, please contact us:
👉 Email: info@hth-inc.com
👉 Website: https://hth-inc.com/business
🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.bam-asr-conversational
All Bambara ASR Dataset
This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-conversational.korean-conversational-speech-corpus
ASR-KCSC: A Korean Conversational Speech Corpus
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Speech Corpus
Language: Korean
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Recording Environment: Indoor
Dataset Description
This open-source dataset consists of 5.22 hours of transcribed Korean conversational speech on certain topics, where 22 conversations between seven pairs of speakers… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/korean-conversational-speech-corpus.ePark_sheng_huo_hui_hua_pian_daily_conversation
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation.personaplex-distill-conversations
PersonaPlex Distillation Conversation Dataset
Teacher-generated multi-turn conversation data for distilling/pruning NVIDIA PersonaPlex 7B
(a Moshi-architecture full-duplex speech-to-speech model).
What this is
Each sample is a real conversation rendered by the PersonaPlex teacher itself:
The student's turns are scripted (generated by Qwen3-8B) and voiced with Piper TTS
The teacher's (PersonaPlex's) responses are improvised live by the model — its real… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/personaplex-distill-conversations.multi-stream-spontaneous-conversation-training-datasets_english
Multi-stream Spontaneous Conversation Training Datasets_English
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Corpus
Language: English
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Dataset Description
The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_english.Miracle-ConversationThis is the dataset curated from ChatGPT with personalized prompt from Our EMNLP2023-findings Miracle
We offer three personality aspects:
'a' = attitude (positive/negative)
'l' = language style (lyrical/plain)
'm' = mental characteristics (critical/emotional)
Shanghai_Dialect_Conversational_Speech_Corpus
Corpus
This dataset is built from Magicdata ASR-CZDIACSC: A CHINESE SHANGHAI DIALECT CONVERSATIONAL SPEECH CORPUS
This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information.
Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Shanghai_Dialect_Conversational_Speech_Corpus.french-conversation+15 hours of speech data from TTS and text file recording.
+9k utterances from various sources, novels, parliamentary debates, professional language.
1000h-us-english-smartphone-conversation
📚 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings)
This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for:
Automatic Speech Recognition (ASR)
Speaker Identification and Gender/Age Analysis
Dialect and Accent Modeling
Multi-speaker Speech Separation
🧾 Dataset Contents
The dataset includes:
metadata.CSV: Metadata including speaker gender, age… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.en-everyday-conversation-asr
English Everyday-Conversation ASR dataset
34777 clips · 39.996 hours · 16 kHz mono · CC-BY-SA 4.0
Assembled from two commercial-safe (CC-BY-SA 4.0) open corpora: EdAcc (spontaneous
dyadic conversation, 11485 clips) and DailyTalk (scripted
everyday-life dialogue, 23292 clips). Splits
({'train': 31297, 'dev': 3480}) are conversation-disjoint (no speaker leakage). Per-row metadata
(gender, accent, style, l1, ...) is preserved so the set can be re-balanced.
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/bbdontcry/en-everyday-conversation-asr.dhivehi-conversations-turn
Dhivehi Conversations (Turn-Based)
This is an experimental synthetic dataset of turn-based Dhivehi conversations created for testing and fine-tuning dialogue models, text-to-speech (TTS), and multi-turn speaker-aware systems.
This dataset is artificially constructed and not based on real conversations. It is intended for research experimentation only and may not always produce contextually accurate results.
Dataset Source
Derived from alakxender/voice-synthetic… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-conversations-turn.English_Natural_Conversation_ASR_STT
BoxlyX English Natural Conversation Sample Dataset (ASR/STT)
📌 Overview
This repository contains high-fidelity, studio-recorded English natural conversation samples designed for training and benchmarking advanced Automatic Speech Recognition (ASR) and Speech-to-Text (STT) models.
This dataset is a curated public sample provided by BoxlyX AI Solution, showcasing our end-to-end capabilities in premium audio data generation, multi-speaker recording environment… See the full description on the dataset page: https://huggingface.co/datasets/BoxlyX/English_Natural_Conversation_ASR_STT.mega-asr-conversational-overlap
Mega-ASR Conversational Overlap
Mega-ASR Conversational Overlap is a deterministic English ASR diagnostic set
derived from AirCaps/mega-asr-noise-a5sv2,
which in turn is sampled from the Mega-ASR training corpus
zhifeixie/Voices-in-the-Wild-2M.
The existing AirCaps dataset evaluates single-utterance acoustic robustness.
This companion dataset evaluates a different failure mode: two-turn conversational
continuity with slight overlap and unequal turn loudness. It does not replace… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-conversational-overlap.malay-conversational-speech-corpus
malay-conversational-speech-corpus
Mirror for https://magichub.com/datasets/malay-conversational-speech-corpus/, license is Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License
Changsha_Dialect_Conversational_Speech_Corpus
Corpus
This dataset is built from Magicdata ASR-CCHSHDIACSC: A CHINESE CHANGSHA DIALECT CONVERSATIONAL SPEECH CORPUS
This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information.
Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed.… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Changsha_Dialect_Conversational_Speech_Corpus.Nanchang_Dialect_Conversational_Speech_Corpus
Corpus
This dataset is built from Magicdata ASR-CNANDIACSC: A CHINESE NANCHANG DIALECT CONVERSATIONAL SPEECH CORPUS
This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information.
Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Nanchang_Dialect_Conversational_Speech_Corpus.dailytalk-conversations-grouped
Dataset Card for "dailytalk-conversations-grouped"
This dataset is intended for testing fine-tuning of Sesame’s CSM-1B (Conversational Speech Model). It is extracted from the DailyTalk dataset.
conversation-bench
Conversation Bench
75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.conversational_englishhuman-robot-conversation-russian
Human-Robot Dataset
The dataset comprises 660+ hours of Russian speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-russian.conversational-s2s-v1
Conversational S2S v1 — dialogues parlés FR + EN
Corpus synthétique de conversations orales multi-tours entre un utilisateur
et un assistant, en français et en anglais, destiné au finetuning
speech-to-speech de LFM2.5-Audio (Liquid AI). Chaque tour est un clip audio
séparé, aligné avec son texte ; les dialogues sont conçus pour être rejoués tour
à tour (user → assistant → user → …).
Projet tts-model-exploration. Produit par
voxtral_datagen_pipeline (s2s-skeletons → remplissage… See the full description on the dataset page: https://huggingface.co/datasets/Rcarvalo/conversational-s2s-v1.Zhengzhou_Dialect_Conversational_Speech_Corpus
Corpus
This dataset is built from Magicdata ASR-CZDIACSC: A CHINESE ZHENGZHOU DIALECT CONVERSATIONAL SPEECH CORPUS
This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information.
Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Zhengzhou_Dialect_Conversational_Speech_Corpus.
