CoolFace
Datasetpublic

allandclive/UgandaLex2

UgandaLex2: A Parallel Text Translation Corpus in 24 Ugandan Languages (3 added languages) UgandaLex Parallel Texts in Ugandan Languages is a remarkable dataset consisting of parallel texts sourced from Bible translations across 21 Ugandan languages. This expansive corpus provides an invaluable resource for studying and analyzing the linguistic variations and nuances within Uganda's diverse language landscape. With aligned texts from various Bible translations, researchers… See the full description on the dataset page: https://huggingface.co/datasets/allandclive/UgandaLex2.

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes25downloads
Dataset Card

UgandaLex2: A Parallel Text Translation Corpus in 24 Ugandan Languages (3 added languages)

UgandaLex Parallel Texts in Ugandan Languages is a remarkable dataset consisting of parallel texts sourced from Bible translations across 21 Ugandan languages. This expansive corpus provides an invaluable resource for studying and analyzing the linguistic variations and nuances within Uganda's diverse language landscape. With aligned texts from various Bible translations, researchers, linguists, and developers can delve into the intricacies of Ugandan languages, explore translation patterns, and investigate the cultural and linguistic heritage of different communities. UgandaLex opens up avenues for advancing research in computational linguistics, cross-linguistic analysis, and the development of language technologies tailored specifically for Ugandan languages.

Languages

Kebu, Acholi, Saamya-Gwe, **Nyoro, Alur, Aringa, Ateso, Ganda, Gwere, Jopadhola, Kakwa, Kinyarwanda, Kumam, Lango, Lugbara, Masaaba, Ng'akarimojong, Nyankore, Nyole, Soga, Swahili, English, Gungu, Keliko, Talinga-Bwisi

Contributors

@allandclive & @oumo_os