datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Adaption-low-resource-audio
Adaption Low-Resource Audio
A low-resource-language subset of
Reubencf/PolyglotAudio,
remastered with Adaption's Adaptive Data
platform. Each row carries the original Tatoeba-derived audio clip
alongside sharpened enhanced_prompt / enhanced_completion columns
so the data is ready for speech-model fine-tuning and evaluation on
languages that are typically under-represented in open ASR/TTS corpora.
Dataset size
3,704 rows of paired audio + text, spanning 10 languages… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-audio.VietMuong-LowResourceLow_Resource_Arabic_Adaptation
