datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UgandaLex
UgandaLex: A Parallel Text Translation Corpus in 21 Ugandan Languages
UgandaLex Parallel Texts in Ugandan Languages is a remarkable dataset consisting of parallel texts sourced from Bible translations across 21 Ugandan languages. This expansive corpus provides an invaluable resource for studying and analyzing the linguistic variations and nuances within Uganda's diverse language landscape. With aligned texts from various Bible translations, researchers, linguists, and developers can… See the full description on the dataset page: https://huggingface.co/datasets/allandclive/UgandaLex.UgandaLex2
UgandaLex2: A Parallel Text Translation Corpus in 24 Ugandan Languages (3 added languages)
UgandaLex Parallel Texts in Ugandan Languages is a remarkable dataset consisting of parallel texts sourced from Bible translations across 21 Ugandan languages. This expansive corpus provides an invaluable resource for studying and analyzing the linguistic variations and nuances within Uganda's diverse language landscape. With aligned texts from various Bible translations, researchers… See the full description on the dataset page: https://huggingface.co/datasets/allandclive/UgandaLex2.
