datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-casual-conversational-speech-golden-dataset-preview
Japanese Casual Conversational Speech Golden Dataset (Preview)
💼 Commercial License & Full Access
This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.
To purchase the full dataset, please contact us:
👉 Email: info@hth-inc.com
👉 Website: https://hth-inc.com/business
🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.bam-asr-conversational
All Bambara ASR Dataset
This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-conversational.mega-asr-conversational-overlap
Mega-ASR Conversational Overlap
Mega-ASR Conversational Overlap is a deterministic English ASR diagnostic set
derived from AirCaps/mega-asr-noise-a5sv2,
which in turn is sampled from the Mega-ASR training corpus
zhifeixie/Voices-in-the-Wild-2M.
The existing AirCaps dataset evaluates single-utterance acoustic robustness.
This companion dataset evaluates a different failure mode: two-turn conversational
continuity with slight overlap and unequal turn loudness. It does not replace… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-conversational-overlap.conversational-speech-dataset
🎙️ Silencio Network: Conversational Speech Dataset
Overview
Sample conversational speech data from Silencio Network's crowdsourced voice AI platform. This dataset contains multi-speaker meeting recordings with word-level transcripts, speaker diarization, and rich demographic metadata.
Each row represents one participant in a meeting and includes 3 audio files:
Audio Column
Description
Format
file_name (speaker audio)
Individual participant's… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/conversational-speech-dataset.realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.tts-conversational-voice-20000h
TTS Voice Dataset
20,000 hours of high-fidelity 48kHz conversational audio across 30+ global, regional, and underrepresented languages, built for text-to-speech, voice cloning, and multilingual speech AI.
This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real audio samples.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/tts-conversational-voice-20000h.conversational_kannada_stt
Conversational Kannada STT
This dataset contains corrected transcriptions of conversational Kannada speech,prepared specifically for fine-tuning Whisper models on conversational and dialectal Kannada.
Unlike many ASR datasets, this release provides pre-computed Whisper input features (log-Mel spectrograms)so you can train/fine-tune Whisper models without raw audio processing.
Dataset Creation
Source Audio: Publicly available YouTube videos in Kannada.
Initial… See the full description on the dataset page: https://huggingface.co/datasets/ShimogaAIteam/conversational_kannada_stt.Global-Conversational-Speech
Global Conversational Speech Dataset
305 hours. 18 locales. Real conversations.
Not scraped from YouTube. Not recorded by anonymous crowds who don't speak the language. Every conversation in this dataset traces back to verified native speakers we know by name.
The [Human] Standard
Most speech datasets are built the same way: scrape the internet, hire anonymous contractors, run it through automated QC, ship it. The result? Models that are confidently wrong.
We… See the full description on the dataset page: https://huggingface.co/datasets/UsergyAI/Global-Conversational-Speech.tts-conversational-voice-20000h
TTS Voice Dataset
20,000 hours of high-fidelity 48kHz conversational audio across 30+ global, regional, and underrepresented languages, built for text-to-speech, voice cloning, and multilingual speech AI.
This repository is a specification and preview listing. The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real audio samples.
Overview
The TTS Voice Dataset is a 20,000-hour collection of… See the full description on the dataset page: https://huggingface.co/datasets/DatoricAI/tts-conversational-voice-20000h.AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE
AMERICAN ENGLISH TRANSCRIBED HI-FI FULL-DUPLEX TWO-SPEAKER CONVERSATIONAL DATASET — SAMPLE
Overview
This open sample from Ocular AI contains four American English conversations between two people, with a separate audio track for each speaker and verbatim transcripts containing segment- and word-level timestamps.
The recordings capture conversational exchanges: repetitions, fillers, false starts, pauses, laughter, and audible breaths. Some conversations begin with… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE.
