datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.Hausa-Synthetic-ASR-Dataset-XTTSSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model.
Sample rate: 24kHz.
Total duration: 574 hours.
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.Synthetic_Turkish_TTS_Data
Synthetic Turkish TTS Data
This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech.
These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/Synthetic_Turkish_TTS_Data.synthetic-speaker-diarization-dataset-fa-large-3000burmese-synthetic-speech-corpus
Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus)
Overview
The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language.
Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-synthetic-speech-corpus.Synthetic_Turkish_TTS_Data
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti Anilosan15 tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: Anilosan15/Synthetic_Turkish_TTS_Data
🔗 Derleyen Platform: VeriPazarı
Sentetik Türkçe TTS Veri Seti (Synthetic Turkish TTS Data)
Bu veri seti, çoklu konuşma senaryoları üzerinden sentetik Türkçe metinler… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Synthetic_Turkish_TTS_Data.Chichewa-Synthetic-ASR-DatasetSynthetic Chichewa ASR dataset generated using a fine-tuned version of the YourTTS model.
Sample rate: 24kHz.
Total duration: 550 hours.
Synthetic-Egy-Speech-Dataset
Synthetic Egyptian Speech Dataset
A curated dataset of 1000 Egyptian Arabic speech samples — the best audio selected across 4 TTS models for each prompt, with transcription text and quality metadata.
Dataset Description
Each entry contains:
id: Unique prompt identifier (e.g., egy_0001)
text: Egyptian Arabic transcription text
audio_path: Path to the best-selected .wav audio file
model: TTS model that produced the best audio (lahgtna, chatterbox_egyptian, egtts_v01, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Synthetic-Egy-Speech-Dataset.Luo-Synthetic-ASR-DatasetSynthetic Dholuo ASR dataset generated using a fine-tuned version of the YourTTS model.
Sample rate: 24kHz.
Total duration: 775 hours.
Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/niamhtracey1/Synthetic-Medical-Speech-Dataset.Hausa-Synthetic-ASR-Dataset-YourTTSSynthetic Hausa ASR dataset generated using a fine-tuned version of the YourTTS model.
Sample rate: 24kHz.
Total duration: 993 hours.
Synthetic_Speech_Data_Project
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LisanneH/Synthetic_Speech_Data_Project.combined_synthetic_datasets_eng_hin_engandhincodemix
Combined Synthetic Datasets (English, Hindi, Code-Mix)
Public ASR training data combining YouTube podcast VAD clips, English/Hinglish podcasts, and synthetic Hinglish entity-normalization speech.
Subsets
Config
Rows
Description
yt_video_transcript
4,100
Hindi-dominant YouTube podcast segments (VAD chunks)
vad_english
1,160
English podcast segments
vad_hindi_english
787
Hindi–English code-mixed podcast segments
synthetic_voice_stt
24,459
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/combined_synthetic_datasets_eng_hin_engandhincodemix.
