datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Adaption-multilingual-speech
This dataset is a remastered version of
Reubencf/multilingual-synthetic-tts
prepared using Adaption's Adaptive Data platform.
Multilingual Speech (Adaption)
10,274 audio + text rows selected from the original 68,677-clip
multilingual synthetic speech corpus, with Adaption-sharpened
enhanced_prompt and enhanced_completion columns. Every row carries
the synthesised audio, the ground-truth text, and language/style/voice
metadata — ready for speech SFT.
Original dataset (for… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-speech.Adaption-low-resource-audio
Adaption Low-Resource Audio
A low-resource-language subset of
Reubencf/PolyglotAudio,
remastered with Adaption's Adaptive Data
platform. Each row carries the original Tatoeba-derived audio clip
alongside sharpened enhanced_prompt / enhanced_completion columns
so the data is ready for speech-model fine-tuning and evaluation on
languages that are typically under-represented in open ASR/TTS corpora.
Dataset size
3,704 rows of paired audio + text, spanning 10 languages… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-audio.
