african-languages
african_languages_translationafrican-languages-corpus
VelkroLM African Languages Corpus
This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative.
This publication is… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-corpus.african-languages-hplt-filtered
VelkroLM African Languages Corpus
This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative.
This publication is… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered.temp_africaNLP_keyword_spotting_for_african_languagesThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.open_math_instruct_v2_translated_african_languagesThis is a set of 41k nvidia/OpenMathInstruct-2 questions translated into 9 African languages using Azure/GPT-4o.
We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language.
LLaVA-Pretrain-558k-10-African-Languages
LLaVA-Pretrain 558K — Translated into 10 African Languages
Machine translation of the LLaVA-Pretrain
caption dataset (blip_laion_cc_sbu_558k, 558,128 image–caption pairs) into 10 low-resource
African languages, for the feature-alignment / pretraining stage of LLaVA-style multimodal SFT.
Languages
Tumbuka (tum), Sepedi / Northern Sotho (nso), Twi / Akan (tw), Chichewa / Nyanja (ny),
Igbo (ig), Nigerian Pidgin (pcm), Moroccan Arabic / Darija (ary), Xhosa (xh)… See the full description on the dataset page: https://huggingface.co/datasets/ketanmore/LLaVA-Pretrain-558k-10-African-Languages.
