datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-wakewordsvoice-light-synthetic-audio
Voice-Light Synthetic Audio
English-only synthetic conversational speech for training and evaluating streaming
turn-taking models. The corpus focuses on end-of-turn prediction, continuation holds,
short backchannels, interruptions, and response timing.
The dataset contains user-side FLAC speech units plus typed conversation plans,
rendering provenance, quality ledgers, and deterministic reconstruction metadata.
Assistant speech is represented as a time-varying… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-synthetic-audio.synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
synthetic_dem
Dataset Card for synthetic_dem
Dataset Summary
The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC).
It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.Synthetic-User-Turn-TTS
Synthetic Malaysian Telco Call-Centre Speech
Synthetic Malaysian call-centre customer utterances, as text and as speech.
The text is fully synthetic dialogue styled after real Malaysian ISP/telco ("Unifi")
call-centre recordings, containing no real customer data. The audio subsets take customer
(user) turns and voice them with a voice-conversion model, keeping only clips an ASR
round-trip confirms are accurate.
Subsets
subset
rows
content
default
4,260… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Synthetic-User-Turn-TTS.capes_synthetic_audio_filteredserena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.synthetic-wakeword-hey_computer
synthetic-wakeword-hey_computer
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey computer".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.german-sohee-synthetic-tts-24k
Flevi Restiti Vici
Bis uns der Weg nur noch nach vorn blieb
This repository contains a synthetic German audiobook and text-to-speech training dataset based on the original novel:
Flevi Restiti ViciBis uns der Weg nur noch nach vorn blieb
The novel was written in German by Maurice Hartmann, who is the author and copyright holder.
Work in progress
This dataset is a work in progress. Future revisions may include corrected transcripts, regenerated audio… See the full description on the dataset page: https://huggingface.co/datasets/Muckylixx/german-sohee-synthetic-tts-24k.synthetic-speech-indicsynthetic-wakewords
synthetic-wakewords
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "multiple wake words".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/synthetic-wakewords.LFM-Audio-IFEval-Synthetic
LFM-Audio IFEval Synthetic — targeted continuation
This release adds 532 verified training tasks, with new validation and test scenarios, to the original 479-task corpus. It targets omitted placeholders/highlights/keywords, exact endings, response and paragraph structure, case/count constraints, and high-count failures. Every target passed mechanical checks, a DeepSeek quality review, Echo synthesis, Whisper large-v3 transcription, and spoken-instruction preservation checks.… See the full description on the dataset page: https://huggingface.co/datasets/Omni-Post-Train/LFM-Audio-IFEval-Synthetic.multivoice-synthetic-speech
Synthetic Voice Samples · Africa
Synthetic speech. No human speaker was recorded for any clip in this dataset.
Generated with afrispeech-synth: text from
africa-corpus, normalised to a
universal orthography with africa-g2p, spoken by
Google Gemini's Live API.
17,010 clips · 38.8 hours · 566 languages · 30 voices
Every clip is a distinct sentence — no sentence is repeated
Each language is read by up to 30 different voices, one sentence per voice
~1.29 hours per voice… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/multivoice-synthetic-speech.cv_mls_psfb_zero_syntheticData used to reproduce all of the experiments in the paper ; https://ieeexplore.ieee.org/document/10720758/
synthetic_transcript_pt
Portuguese Speech Dataset with Multiple Training Configurations
A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms.
🎯 Dataset Configurations Overview
This dataset provides three carefully curated subsets to enable comprehensive speech recognition research:
Configuration
Training Data
Validation
Test
Total Samples
Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.slurp_synthetic_barkSynthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.masri_synthetic
Dataset Card for masri_synthetic
Dataset Summary
The MASRI-SYNTHETIC is a corpus made out of synthesized speech in Maltese. The text-to-speech (TTS) system utilized to produce the utterances was developed by the Research & Development Department of Crimsonwing p.l.c.
The sentences used to create the corpus were extracted from the MLRS Corpus, which is a corpus of written or transcribed Maltese divided into different genres, including: culture, news, academic, religion… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/masri_synthetic.cv_for_spd_fr_syntheticbengali-diarization-synthetic-v3echo-synthetic-diarization
Echo (Synthetic Set for Diarization)
This dataset is a synthetic dataset generated for evaluation of speaker
diarization models. It contains approximately two hours of speech data, each
file 60 seconds long, with and without overlap, with 2--5 speakers per file.
This dataset has been built using Echo.
romanian_speech_dataset_with_15_percent_6_speakers_synthetic_datanazrah-synthetic-datasethi_luna_synthetic_audio_v1300k audio files synthetically generated by VITS using https://github.com/dscripka/synthetic_speech_dataset_generation?tab=readme-ov-file
Command used
python generate_clips.py \
--model VITS \
--enable_gpu \
--text "Hey, Luna" \
--N 300000 \
--max_per_speaker 1 \
--output_dir /luna_audio
synthetic_speech_commands_PA_taggedmixture_ami_synthetic_bigsynthetic_transcript_nl
Dutch Synthetic Speech Transcripts
This dataset contains 34,898 synthetic Dutch speech samples generated using GPT-4o-mini for transcript creation and OpenAI's TTS-1 model for speech synthesis. It was designed to augment Automatic Speech Recognition (ASR) training for low-resource scenarios, matching the linguistic distribution of Common Voice 17.0 Dutch.
Dataset Description
Purpose
This dataset addresses the challenge of limited labeled speech data for Dutch… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_nl.syntheticarabic-english-code-switching-synthetic-asr
Synthetic Arabic-English Code-Switched Speech for ASR
This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository.
Configurations
Configuration
Train
Test
Publication status
synthetic
8,655
962
Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.saudi-tts-synthetic-200k
