datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acl-6060
ACL 60/60
Dataset details
ACL 60/60 evaluation sets for multilingual translation of ACL 2022 technical presentations into 10 target languages.
Citation
@inproceedings{salesky-etal-2023-evaluating,
title = "Evaluating Multilingual Speech Translation under Realistic Conditions with Resegmentation and Terminology",
author = "Salesky, Elizabeth and
Darwish, Kareem and
Al-Badrashiny, Mohamed and
Diab, Mona and
Niehues, Jan"… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/acl-6060.Sukuma-Voices-ACL
Sukuma Voices Dataset 🎙️
The first publicly available speech corpus for Sukuma (Kisukuma), a Bantu language spoken by approximately 10 million people in northern Tanzania. This dataset supports speech-to-text, text-to-speech, and speech evaluation tasks.
Dataset Description
Sukuma Voices addresses the critical gap in speech technology resources for one of Africa's most severely under-resourced languages. The dataset includes both human recordings and… See the full description on the dataset page: https://huggingface.co/datasets/sartifyllc/Sukuma-Voices-ACL.
