datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SynthGT
SynthGT
A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment
Authors
Silas Antonisen, Iván López-Espejo
Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing.
Overview
SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations.
The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.barranquenho-ipa-dict-synthetic
Barranquenho IPA Pronunciation Dictionary
The first and only IPA pronunciation dictionary of Barranquenho — the
Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal),
a mixed system born of centuries of Portuguese–Spanish (Extremaduran /
Andalusian) contact on the raia. Every headword is written in the
Convenção Ortográfica do Barranquenho (2025) orthography and paired with a
broad-phonemic IPA transcription plus Portuguese and Spanish glosses.
This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.BLURB-synth
BLURB-synth: Synthetic audio data based on BLURB corpora
Dataset Summary
Synthetic audio data based on BLURB corpora. More details coming soon...
Supported Tasks and Leaderboards
Biomedical Language Understanding and Reasoning Benchmark (BLURB)
Text-to-Speech
Automatic-Speech-Recognition
Languages
English
Data Structure
Data Instances
Coming soon...
Data Fields
Coming soon...… See the full description on the dataset page: https://huggingface.co/datasets/uy-rrodriguez/BLURB-synth.somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.
