datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yoruba_text_c3Yoruba Text C3 is the largest Yoruba texts collected and used to train FastText embeddings in the
YorubaTwi Embedding paper: https://www.aclweb.org/anthology/2020.lrec-1.335/yoruba-normalization-pairs
Normalization pairs dataset
What this is
24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code.
The library
This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.Code-170k-yoruba
Dataset Description
Code-170k-yoruba is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Yoruba, making coding education accessible to Yoruba speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Yoruba language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-yoruba.task612_yorubabbc_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task612_yorubabbc_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task612_yorubabbc_classification.yoruba-cultural-reasoning-blindspots# Yoruba Cultural Reasoning Blind Spots in Frontier Models
## Overview
This dataset captures the "blind spots" of frontier base models when evaluating non-Western abstract reasoning, specifically focusing on Yoruba proverbs. It contains 10 diverse examples demonstrating "generative collapse" and "Cultural Hallucination."
## 1. Loading the Model
This evaluation was conducted using a standard Google Colab environment with a T4 GPU. The model evaluated was `Qwen/Qwen2.5-1.5B`. It was loaded in… See the full description on the dataset page: https://huggingface.co/datasets/saaga/yoruba-cultural-reasoning-blindspots.adtc-agri-yoruba-training-mix
ADTC Agri-Yoruba Training Mix
An English/Yoruba instruction tuning dataset for agricultural extension advisory, built for the Africa Deep Tech Challenge 2026 Laptop LLM track by Team Support Vector. Used to fine tune llama-3.2-3b-agri-yoruba.
Composition
43,640 examples total, in chat message format (system, user, assistant):
Source
Rows
Share
naija_yoruba
27,464
63%
ai71_agrillm
14,996
34%
identity_control
800
2%
behavior_control
240
1%… See the full description on the dataset page: https://huggingface.co/datasets/1nnocent/adtc-agri-yoruba-training-mix.foundry-y-yoruba-corpus
Yorùbá News & Tasks — a contamination-proof, 100% human Yorùbá corpus
A 19,997-row supervised fine-tuning corpus ({"prompt","completion"} JSONL) for
adapting a language model to Yorùbá, built for the Adaption AutoScientist
challenge (language track). Every row is human-written, from an approved,
train-split-only source. No synthetic or LLM-generated text.
SHA-256: a711692096719c0d11f8e9c3784061211733029bf32a20e36b7e920488f8cc0c
Rows: 19,997 · Format: JSONL, prompt + completion… See the full description on the dataset page: https://huggingface.co/datasets/Enochid/foundry-y-yoruba-corpus.yoruba_pidgin_agriculture_data
