datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
syntheory
Dataset Card for SynTheory
Dataset Summary
SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions).
Each of these 7 concepts has its own config.
tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time).
time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.synthetic-wakewordsvoice-light-synthetic-audio
Voice-Light Synthetic Audio
English-only synthetic conversational speech for training and evaluating streaming
turn-taking models. The corpus focuses on end-of-turn prediction, continuation holds,
short backchannels, interruptions, and response timing.
The dataset contains user-side FLAC speech units plus typed conversation plans,
rendering provenance, quality ledgers, and deterministic reconstruction metadata.
Assistant speech is represented as a time-varying… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-synthetic-audio.synthetic_dem
Dataset Card for synthetic_dem
Dataset Summary
The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC).
It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
fluent_speech_commands_synth
Dataset Card for "fluent_speech_commands_synth"
More Information needed
Synthetic-User-Turn-TTS
Synthetic Malaysian Telco Call-Centre Speech
Synthetic Malaysian call-centre customer utterances, as text and as speech.
The text is fully synthetic dialogue styled after real Malaysian ISP/telco ("Unifi")
call-centre recordings, containing no real customer data. The audio subsets take customer
(user) turns and voice them with a voice-conversion model, keeping only clips an ASR
round-trip confirms are accurate.
Subsets
subset
rows
content
default
4,260… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Synthetic-User-Turn-TTS.SynthGT
SynthGT
A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment
Authors
Silas Antonisen, Iván López-Espejo
Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing.
Overview
SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations.
The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.librispeech_synth
Dataset Card for "librispeech_synth"
More Information needed
pony-speechicelandic_asr
Icelandic ASR Collection
This repository collects six Icelandic speech corpora in directly loadable
Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a
convenience repackaging: the linked CLARIN-IS records and original dataset
repositories remain the canonical sources and should be cited when using the
data.
No configuration is selected by default. Choose a corpus configuration and,
for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.capes_synthetic_audio_filteredvocal_imitation_synth
Dataset Card for "vocal_imitation_synth"
More Information needed
crema_d_synth
Dataset Card for "crema_d_synth"
More Information needed
voxceleb1_synthmaestro_synth
Dataset Card for "maestro_synth"
More Information needed
librispeech_asr_test_48k_synthtorgo_synthvox_lingua_top10_synthsyntheory_plus
Notes
Dataset Viewer is disabled as we wanted to keep everything in WAV format with CSV metadata instead of Parquet files and the dataset is too large for Dataset Viewer to index properly.
Dataset Authors
Derek Kwan and Patrick Donnelly
Related Paper
The paper that introduces this dataset is "Probing for Advanced Music Theory Concepts in Generative Music Models" by Derek Kwan and Patrick Donnelly presented at EvoMUSART 2026
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/derekxkwan/syntheory_plus.vocalset_synthspeech_accent_archive_synthlibrispeech_asr_test_synthvocalset_synth
Dataset Card for "vocalset_synth"
More Information needed
opensinger_synthserena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.voxceleb1_synthsynthetic-wakeword-hey_computer
synthetic-wakeword-hey_computer
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey computer".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.synthetic-speech-indiceasycall_synth
