datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.romanian_speech_dataset_with_15_percent_6_speakers_synthetic_datanazrah-synthetic-datasetHausa-Synthetic-ASR-Dataset-XTTSSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model.
Sample rate: 24kHz.
Total duration: 574 hours.
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.synthetic-gujarati-tts-datasetSynthetic_Turkish_TTS_Data
Synthetic Turkish TTS Data
This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech.
These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/Synthetic_Turkish_TTS_Data.romanian_speech_dataset_with_20_percent_4_speakers_synthetic_datasynthetic-speaker-diarization-dataset-fa-large-3000burmese-synthetic-speech-corpus
Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus)
Overview
The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language.
Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-synthetic-speech-corpus.synthetic-speaker-diarization-dataset-hindisynthetic-data-japanfr_en_synthetic_audio_datasetsynthetic-data-indonesia_testsynthetic-speaker-diarization-datasetsynthetic-data-indonesiasynthetic-speaker-diarization-dataset-hindi-largeromanian_speech_dataset_with_40_percent_8_speakers_synthetic_datasynthetic-data-indonesia_2_4synthetic_dataset_jpn_2_more_speakerssynthetic-data-indonesia_2_4_updatedSynthetic_Turkish_TTS_Data
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti Anilosan15 tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: Anilosan15/Synthetic_Turkish_TTS_Data
🔗 Derleyen Platform: VeriPazarı
Sentetik Türkçe TTS Veri Seti (Synthetic Turkish TTS Data)
Bu veri seti, çoklu konuşma senaryoları üzerinden sentetik Türkçe metinler… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Synthetic_Turkish_TTS_Data.Emotion_conversation_synthetic_datasetsynthetic_dataset_jpnsynthetic_dataset_jpn_volumessynthetic-speaker-diarization-datasetChichewa-Synthetic-ASR-DatasetSynthetic Chichewa ASR dataset generated using a fine-tuned version of the YourTTS model.
Sample rate: 24kHz.
Total duration: 550 hours.
synthetic-speaker-diarization-dataset-fasynthetic dataset generated from persian common voice.
Emotion_conversation_synthetic_dataset_shortSynthetic-Egy-Speech-Dataset
Synthetic Egyptian Speech Dataset
A curated dataset of 1000 Egyptian Arabic speech samples — the best audio selected across 4 TTS models for each prompt, with transcription text and quality metadata.
Dataset Description
Each entry contains:
id: Unique prompt identifier (e.g., egy_0001)
text: Egyptian Arabic transcription text
audio_path: Path to the best-selected .wav audio file
model: TTS model that produced the best audio (lahgtna, chatterbox_egyptian, egtts_v01, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Synthetic-Egy-Speech-Dataset.
