datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
african_languages_translationafrican-languages-corpus
VelkroLM African Languages Corpus
This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative.
This publication is… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-corpus.african-languages-hplt-filtered
VelkroLM African Languages Corpus
This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative.
This publication is… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered.temp_africaNLP_keyword_spotting_for_african_languagesThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.open_math_instruct_v2_translated_african_languagesThis is a set of 41k nvidia/OpenMathInstruct-2 questions translated into 9 African languages using Azure/GPT-4o.
We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language.
LLaVA-Pretrain-558k-10-African-Languages
LLaVA-Pretrain 558K — Translated into 10 African Languages
Machine translation of the LLaVA-Pretrain
caption dataset (blip_laion_cc_sbu_558k, 558,128 image–caption pairs) into 10 low-resource
African languages, for the feature-alignment / pretraining stage of LLaVA-style multimodal SFT.
Languages
Tumbuka (tum), Sepedi / Northern Sotho (nso), Twi / Akan (tw), Chichewa / Nyanja (ny),
Igbo (ig), Nigerian Pidgin (pcm), Moroccan Arabic / Darija (ary), Xhosa (xh)… See the full description on the dataset page: https://huggingface.co/datasets/ketanmore/LLaVA-Pretrain-558k-10-African-Languages.african-languages-filtered
Filtered African-language datasets
This repository contains non-destructive filtered derivatives of publicly accessible Hugging Face datasets relevant to Hausa, Nigerian languages, and selected African languages. The original repositories remain the authoritative sources and were not modified.
Scope and provenance
Each JSONL file preserves the source repository, source split, and source row index in _source_repo, _source_split, and _source_row_index.… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-filtered.african-languages-speech
WAXAL African-language speech — filtered wave 1
This repository contains bounded, filtered WAXAL audio/transcript pairs for Hausa (hau_tts), Yoruba (yor_tts), and Igbo (ibo_tts). Each configuration is split into tar archives containing an audio file plus a JSON record with transcript, language, speaker, source ID, and provenance.
The upstream source is google/WaxalNLP and the WAXAL paper is arXiv:2602.02734. Consult the upstream dataset card for the exact component license and… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-speech.ratman-stopword-lists-for-african-languages
Dataset Card for Stopword Lists for African Languages
Dataset Summary
Context:
Some words, like “the” or “and” in English, are used a lot in speech and writing. For most Natural Language Processing applications, you will want to remove these very frequent words. This is usually done using a list of “stopwords” which has been complied by hand.
Content:
This project uses the source texts provided by the African Storybook Project as a corpus and… See the full description on the dataset page: https://huggingface.co/datasets/chrisjay/ratman-stopword-lists-for-african-languages.african-languages-catalog
VelkroLM African-language corpus catalog
This catalog organizes the current audited African-language publication waves.
Published repositories
Text corpus: https://huggingface.co/datasets/VelkroLM/african-languages-corpus
Personal text mirror: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered
Speech wave 1: https://huggingface.co/datasets/VelkroLM/african-languages-speech
HPLT source: https://hplt-project.org/datasets/v3.0
WAXAL source:… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-catalog.African-Languages_Sentiments
African Languages Sentiment Dataset (Hausa, Yorùbá, Swahili)
A stitched multi-source sentiment classification dataset combining three
independently collected sentiment corpora for Hausa, Yorùbá, and Swahili,
built for the Adaption Labs AutoScientist Challenge
(Language category).
Companion model: fine-tuned weights trained on the adapted version of this dataset via
AutoScientist are released separately at… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/African-Languages_Sentiments.big_math_translated_african_languages
Big Math Translated -- African Languages
This is a set of 41k SynthLabsAI/Big-Math-RL-Verified questions translated into 9 African languages using Azure/GPT-4o.
We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language.
multi-open
African Languages Lab Multi-Open
multi-open is the open-source multilingual subset released by the
African Languages Lab. It contains English-target
parallel text for 31 African languages.
Project website: https://the-african-languages-lab.github.io/
The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African
NLPIssaka et al., ACL 2026.
The paper presents All Lab's broader collaborative program: systematic and quality-controlled
data infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/multi-open.english-south-african-languagesafrican-languages-filtered
Filtered African-language datasets
This repository contains non-destructive filtered derivatives of publicly accessible Hugging Face datasets relevant to Hausa, Nigerian languages, and selected African languages. The original repositories remain the authoritative sources and were not modified.
Scope and provenance
Each JSONL file preserves the source repository, source split, and source row index in _source_repo, _source_split, and _source_row_index.… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-filtered.proxy-mt-translations
Proxy-MT Translations
English→X machine translations generated with vLLM
across 50 open-weight LLMs on three evaluation benchmarks. This dataset holds the
raw model outputs (one CSV per model × dataset × target language); metric scores
(BLEU / chrF / COMET / MetricX) live in proxy-mt-eval-scores.
Layout
flores-200/<model>/eng-<lang>.csv # 119 target languages
ntrex/<model>/eng-<lang>.csv # 87 target languages
wmt24/<model>/eng-<lang>.csv # 51… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-translations.kasagadi
Kasagadi — Ghanaian Radio Broadcast Fact-Check Dataset
This is a multilingual dataset of transcribed, translated, and AI fact-checked segments from live radio broadcasts across Ghana. It is a Ghanaian initiative, covering Twi-language broadcasts from two Ghanaian FM stations and Hausa-language broadcasts from a third Ghanaian FM station serving Ghana's Zongo communities.
Dataset Summary
Station
Language
Broadcasts
Segments
Hours
Date Range
Angel FM
Twi… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/kasagadi.adaption-african-languages-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-african_languages_narratives
This dataset contains narrative texts and folktales written in various African languages, including Igbo, Yoruba, Hausa, and lesser-documented tongues. The samples feature rich cultural storytelling, descriptive prose, and the use of reduplication for emphasis. The content ranges from daily life scenes to mythical encounters, serving as a resource for… See the full description on the dataset page: https://huggingface.co/datasets/ChiamakaNwokolo/adaption-african-languages-narratives.african-languages-natlasproxy-mt-eval-scores
Proxy-MT Eval Scores
Corpus-level MT metrics for 50 open-weight LLMs on the translations in
proxy-mt-translations.
Computed by evaluate_mt.py (BLEU, chrF++, ROUGE-L, METEOR, XCOMET-XL, SSA-COMET).
MetricX is backfilled separately and may still be empty in this snapshot.
Layout
<model>/flores-200.csv
<model>/ntrex.csv
<model>/wmt24.csv
Each CSV has one row per eng-<lang> pair:
column
description
translation-pair
e.g. eng-yor
bleu
sacrebleu corpus BLEU… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-eval-scores.all-lab-text-monoproxy-mt-benchmark-scores
Proxy-MT Benchmark Scores
Multilingual benchmark results for 50 open-weight LLMs, evaluated with the
lm-evaluation-harness via a vLLM
backend. Covers reasoning, comprehension, and knowledge tasks with an emphasis on
African and other lower-resource languages.
Layout
scores/<model>.csv # parsed per-language scores (tidy, ready to plot)
raw/<model>/.../results_*.json # raw lm-eval-harness result files
raw/<model>/raw_log.txt # full evaluation… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-benchmark-scores.qwen-blindspots-african-languages
Blind Spots of Qwen3.5-0.8B
Overview
This dataset documents failure cases of the base modelQwen3.5-0.8B.
The model has approximately 0.8 billion parameters and is designed as a lightweight multilingual language model.
The goal of this dataset is to highlight cases where the model produces incorrect, misleading, or suboptimal outputs across diverse tasks including translation, reasoning, and regional knowledge.
The dataset focuses particularly on low-resource African… See the full description on the dataset page: https://huggingface.co/datasets/ouilyh/qwen-blindspots-african-languages.all-lab-speech
All Lab Speech
Cleaned African-language speech with embedded, playable audio (HF Audio). One config per language (<lang> = transcribed, <lang>_manifest = audio-only); the audio column sits right after audio_id and plays in the dataset viewer. Splits (train/validation/test) come from the source split labels.
from datasets import load_dataset
ds = load_dataset("African-Languages-Lab/all-lab-speech", "afrikaans")
Columns
audio_id, audio (playable), transcript… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/all-lab-speech.afrilion-african-languagesg2p-african-languages
G2P African Languages — Mooré · Bambara · Dioula
Projet PFE : Leveraging Phonemic Features for Cross-lingual NLP in African LanguagesCITADEL, Ouagadougou, Burkina Faso — 2025-2026
Dataset graphème-to-phonème (G2P) pour trois langues africaines sous-dotées du Burkina Faso.
Il couvre le Mooré (famille Gur), le Bambara et le Dioula (famille Mandé), et fournit
deux transcriptions phonémiques par entrée : une avec les marques tonales issues du
dictionnaire source, une sans tons —… See the full description on the dataset page: https://huggingface.co/datasets/Uriath/g2p-african-languages.african_languages_translationall-lab-text-multi
