datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yoruba-cfm-latentsyoruba-normalization-pairs
Normalization pairs dataset
What this is
24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code.
The library
This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.TinyStories_yoruba
TinyStories English-Igbo Parallel Corpus
Description
TBD
Composition
TBD
Usage
TBD
Acknowledgments
TBD
License
The translated datasets are released under Apache2.0, consistent with the original TinyStories dataset's licensing terms. Please refer to Microsoft's official release for further details on the licensing of the TinyStories dataset.
About the Authors
Christopher Ibe and Okezie Okoye continue to lead Hypa AI towards new… See the full description on the dataset page: https://huggingface.co/datasets/ccibeekeoc42/TinyStories_yoruba.English-Yoruba-Dictionaryowe-yoruba-training-datayoruba-proverbs-parallel-corporaParralel corpora for yoruba to english.
Source: http://yoruba.unl.edu/yoruba1.html
DollyHHRLHF_yoruba
DollyHHRLHF English-Yoruba Parallel Corpus
Description
TBD
Composition
TBD
Usage
TBD
Acknowledgments
TBD
License
The translated datasets are released under Apache2.0, consistent with the original TinyStories dataset's licensing terms. Please refer to Microsoft's official release for further details on the licensing of the TinyStories dataset.
About the Authors
Christopher Ibe and Okezie Okoye continue to lead Hypa AI towards new… See the full description on the dataset page: https://huggingface.co/datasets/ccibeekeoc42/DollyHHRLHF_yoruba.Yoruba-linguistics-dataset
Yoruba Qwen Fine-Tuned Model
Overview
This model is a LoRA fine-tuned version of Qwen2.5-0.5B-Instruct developed for Yoruba language instruction-following tasks.
The project explores the use of parameter-efficient fine-tuning for low-resource African language NLP, with a particular focus on Yoruba.
Dataset
The dataset was adapted using Adaption Lab and contains:
Total records: 3,598
Training samples: 3,238
Validation samples: 360
Language: Yoruba… See the full description on the dataset page: https://huggingface.co/datasets/Kolawolettyy12/Yoruba-linguistics-dataset.adtc-agri-yoruba-training-mix
ADTC Agri-Yoruba Training Mix
An English/Yoruba instruction tuning dataset for agricultural extension advisory, built for the Africa Deep Tech Challenge 2026 Laptop LLM track by Team Support Vector. Used to fine tune llama-3.2-3b-agri-yoruba.
Composition
43,640 examples total, in chat message format (system, user, assistant):
Source
Rows
Share
naija_yoruba
27,464
63%
ai71_agrillm
14,996
34%
identity_control
800
2%
behavior_control
240
1%… See the full description on the dataset page: https://huggingface.co/datasets/1nnocent/adtc-agri-yoruba-training-mix.naija-agri-dataset-yoruba-hausafoundry-y-yoruba-corpus
Yorùbá News & Tasks — a contamination-proof, 100% human Yorùbá corpus
A 19,997-row supervised fine-tuning corpus ({"prompt","completion"} JSONL) for
adapting a language model to Yorùbá, built for the Adaption AutoScientist
challenge (language track). Every row is human-written, from an approved,
train-split-only source. No synthetic or LLM-generated text.
SHA-256: a711692096719c0d11f8e9c3784061211733029bf32a20e36b7e920488f8cc0c
Rows: 19,997 · Format: JSONL, prompt + completion… See the full description on the dataset page: https://huggingface.co/datasets/Enochid/foundry-y-yoruba-corpus.adapted-yoruba-sft-autoscientist-part2yoruba-datasold-sbpn-20260825-benchmarkyoruba-bolanle-sbpn-20260825-benchmarkyoruba-second-sbpn-20260826-benchmarkadapted-yoruba-sft-autoscientistadapted-yoruba-sft-autoscientistyoruba_pidgin_agriculture_dataTest_Yorubaadaption-yoruba-conversational-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-yoruba_conversational_pairs
This dataset contains prompt-completion pairs in the Yoruba language, featuring conversational exchanges between a user and an AI assistant. The content covers various interpersonal scenarios including conflict resolution, clarification requests, and assurances of privacy. Each entry demonstrates polite, context-aware responses aimed at facilitating… See the full description on the dataset page: https://huggingface.co/datasets/BOLAJOKO03/adaption-yoruba-conversational-pairs.yoruba-academic-wikipedia-qa
Yorùbá Academic Wikipedia Q&A Dataset
This dataset contains ~5,000 instruction-style academic question–answer pairs
in pure Yorùbá with tone marks, derived from Yorùbá Wikipedia.
Format
Each sample follows:
instruction: academic question in Yorùbá
input: empty
output: academic answer in Yorùbá
Intended Use
Instruction tuning
Low-resource language modeling
Educational assistants
Source
Yorùbá Wikipedia (cleaned & deduplicated)
yoruba-corpus-synth
