datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
luganda-fln-training-data
Luganda FLN Training Data
Training data for foundational literacy and numeracy (FLN) models targeting Ugandan primary school teachers (P1–P3). Designed to train small language models (1B parameters) to generate pedagogically sound content in Luganda and English.
Dataset Description
This dataset contains 1,368 training examples across four complementary splits, each targeting different aspects of teacher pedagogical content knowledge for early literacy instruction.… See the full description on the dataset page: https://huggingface.co/datasets/CraneAILabs/luganda-fln-training-data.uganda-crop-qa-luganda
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
uganda_crop_qa_luganda
This dataset contains question-answer pairs focused on agricultural advice for crops like cassava, bananas, and maize in Uganda, featuring content in both English and Luganda. The samples cover pest control, disease identification, planting schedules, and market information tailored to local farmers. It is structured for question-answering tasks with… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/uganda-crop-qa-luganda.luganda-bilingual-literacy-exercises
Luganda-English Bilingual Literacy Exercises (P1–P3)
3,472 structured bilingual exercises for Ugandan primary school literacy instruction (Primary 1 through Primary 3). Each exercise contains parallel English and Luganda versions with questions, answers, and explanations.
Dataset Description
Grade
Exercises
File
P1
1,157
data/p1_exercises.json
P2
1,135
data/p2_exercises.json
P3
1,180
data/p3_exercises.json
Total
3,472
Exercise Types… See the full description on the dataset page: https://huggingface.co/datasets/CraneAILabs/luganda-bilingual-literacy-exercises.Luganda010224adaption-luganda-english-literacy-p1p3
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-luganda_english_literacy_p1p3
This dataset contains 3,472 structured bilingual literacy exercises designed for Ugandan primary school students in grades P1 through P3. The content focuses on phonics, reading comprehension, and vocabulary building through question-answer pairs in both Luganda and English. Samples include tasks such as identifying vowel sounds, syllable… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-luganda-english-literacy-p1p3.luganda_english_dataset_78Knjogerera_english_luganda_corpusadaption-luganda-fln-teacher-training
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-luganda_fln_teacher_training
This dataset contains bilingual educational resources in Luganda and English focused on foundational literacy and numeracy for primary education in Uganda. It includes multiple-choice questions, translation tasks, and detailed pedagogical explanations covering topics like reading fluency, comprehension strategies, and vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-luganda-fln-teacher-training.luganda-english-parallel-corpus
English-Luganda Parallel Corpus for Translation
Dataset Description
This dataset contains parallel sentences in English (en) and Luganda (lg), designed primarily for training and fine-tuning machine translation models. The data consists of sentence pairs extracted from a source document.
Languages
English (en)
Luganda (lg) - ISO 639-1 code: lg
Data Format
The dataset is provided in a format compatible with the Hugging Face datasets library. Each… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-parallel-corpus.luganda-english-bible-corpus
Bible English-Luganda Parallel Corpus
Dataset Description
This dataset contains 32,291 parallel sentences in English (en) and Luganda (lg), derived from biblical texts. It is designed primarily for training and fine-tuning machine translation models, particularly in low-resource language scenarios.
Languages
English (en)
Luganda (lg) - ISO 639-1 code: lg
Data Format
The dataset is provided in a format compatible with the Hugging Face datasets… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-bible-corpus.
