CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aletheia-ng /african_languages_translationtext1M<n<10M1 likes661 downloads1y agoHugging Face02VelkroLM /african-languages-corpus VelkroLM African Languages Corpus This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative. This publication is… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-corpus.text1M<n<10M0 likes243 downloads1mo agoHugging Face03rufatronics /african-languages-hplt-filtered VelkroLM African Languages Corpus This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative. This publication is… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered.text1M<n<10M0 likes235 downloads1mo agoHugging Face04cdleong /temp_africaNLP_keyword_spotting_for_african_languagesThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.0 likes65 downloads4y agoHugging Face05taresco /open_math_instruct_v2_translated_african_languagesThis is a set of 41k nvidia/OpenMathInstruct-2 questions translated into 9 African languages using Azure/GPT-4o. We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language. text10K<n<100K0 likes46 downloads9mo agoHugging Face06ketanmore /LLaVA-Pretrain-558k-10-African-Languagesgated LLaVA-Pretrain 558K — Translated into 10 African Languages Machine translation of the LLaVA-Pretrain caption dataset (blip_laion_cc_sbu_558k, 558,128 image–caption pairs) into 10 low-resource African languages, for the feature-alignment / pretraining stage of LLaVA-style multimodal SFT. Languages Tumbuka (tum), Sepedi / Northern Sotho (nso), Twi / Akan (tw), Chichewa / Nyanja (ny), Igbo (ig), Nigerian Pidgin (pcm), Moroccan Arabic / Darija (ary), Xhosa (xh)… See the full description on the dataset page: https://huggingface.co/datasets/ketanmore/LLaVA-Pretrain-558k-10-African-Languages.image-to-text1M<n<10M0 likes41 downloads1d agoHugging Face07VelkroLM /african-languages-filtered Filtered African-language datasets This repository contains non-destructive filtered derivatives of publicly accessible Hugging Face datasets relevant to Hausa, Nigerian languages, and selected African languages. The original repositories remain the authoritative sources and were not modified. Scope and provenance Each JSONL file preserves the source repository, source split, and source row index in _source_repo, _source_split, and _source_row_index.… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-filtered.100K<n<1M0 likes41 downloads1mo agoHugging Face08VelkroLM /african-languages-speech WAXAL African-language speech — filtered wave 1 This repository contains bounded, filtered WAXAL audio/transcript pairs for Hausa (hau_tts), Yoruba (yor_tts), and Igbo (ibo_tts). Each configuration is split into tar archives containing an audio file plus a JSON record with transcript, language, speaker, source ID, and provenance. The upstream source is google/WaxalNLP and the WAXAL paper is arXiv:2602.02734. Consult the upstream dataset card for the exact component license and… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-speech.0 likes37 downloads1mo agoHugging Face09chrisjay /ratman-stopword-lists-for-african-languages Dataset Card for Stopword Lists for African Languages Dataset Summary Context: Some words, like “the” or “and” in English, are used a lot in speech and writing. For most Natural Language Processing applications, you will want to remove these very frequent words. This is usually done using a list of “stopwords” which has been complied by hand. Content: This project uses the source texts provided by the African Storybook Project as a corpus and… See the full description on the dataset page: https://huggingface.co/datasets/chrisjay/ratman-stopword-lists-for-african-languages.text1 likes31 downloads4y agoHugging Face10VelkroLM /african-languages-catalog VelkroLM African-language corpus catalog This catalog organizes the current audited African-language publication waves. Published repositories Text corpus: https://huggingface.co/datasets/VelkroLM/african-languages-corpus Personal text mirror: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered Speech wave 1: https://huggingface.co/datasets/VelkroLM/african-languages-speech HPLT source: https://hplt-project.org/datasets/v3.0 WAXAL source:… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-catalog.text0 likes31 downloads1mo agoHugging Face11gospelgit /African-Languages_Sentiments African Languages Sentiment Dataset (Hausa, Yorùbá, Swahili) A stitched multi-source sentiment classification dataset combining three independently collected sentiment corpora for Hausa, Yorùbá, and Swahili, built for the Adaption Labs AutoScientist Challenge (Language category). Companion model: fine-tuned weights trained on the adapted version of this dataset via AutoScientist are released separately at… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/African-Languages_Sentiments.texttext-classification10K<n<100K0 likes30 downloads3mo agoHugging Face12taresco /big_math_translated_african_languages Big Math Translated -- African Languages This is a set of 41k SynthLabsAI/Big-Math-RL-Verified questions translated into 9 African languages using Azure/GPT-4o. We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language. texttext-generation10K<n<100K0 likes29 downloads9mo agoHugging Face13African-Languages-Lab /multi-opengated African Languages Lab Multi-Open multi-open is the open-source multilingual subset released by the African Languages Lab. It contains English-target parallel text for 31 African languages. Project website: https://the-african-languages-lab.github.io/ The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLPIssaka et al., ACL 2026. The paper presents All Lab's broader collaborative program: systematic and quality-controlled data infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/multi-open.tabulartranslation10M<n<100M2 likes27 downloads3mo agoHugging Face14Or4kool /english-south-african-languagestext1M<n<10M0 likes21 downloads4mo agoHugging Face15rufatronics /african-languages-filtered Filtered African-language datasets This repository contains non-destructive filtered derivatives of publicly accessible Hugging Face datasets relevant to Hausa, Nigerian languages, and selected African languages. The original repositories remain the authoritative sources and were not modified. Scope and provenance Each JSONL file preserves the source repository, source split, and source row index in _source_repo, _source_split, and _source_row_index.… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-filtered.100K<n<1M0 likes20 downloads1mo agoHugging Face16African-Languages-Lab /proxy-mt-translationsgated Proxy-MT Translations English→X machine translations generated with vLLM across 50 open-weight LLMs on three evaluation benchmarks. This dataset holds the raw model outputs (one CSV per model × dataset × target language); metric scores (BLEU / chrF / COMET / MetricX) live in proxy-mt-eval-scores. Layout flores-200/<model>/eng-<lang>.csv # 119 target languages ntrex/<model>/eng-<lang>.csv # 87 target languages wmt24/<model>/eng-<lang>.csv # 51… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-translations.texttranslation10M<n<100M0 likes18 downloads1mo agoHugging Face17African-Languages-Lab /kasagadigated Kasagadi — Ghanaian Radio Broadcast Fact-Check Dataset This is a multilingual dataset of transcribed, translated, and AI fact-checked segments from live radio broadcasts across Ghana. It is a Ghanaian initiative, covering Twi-language broadcasts from two Ghanaian FM stations and Hausa-language broadcasts from a third Ghanaian FM station serving Ghana's Zongo communities. Dataset Summary Station Language Broadcasts Segments Hours Date Range Angel FM Twi… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/kasagadi.audiotext-classification100K<n<1M0 likes17 downloads3mo agoHugging Face18ChiamakaNwokolo /adaption-african-languages-narratives This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-african_languages_narratives This dataset contains narrative texts and folktales written in various African languages, including Igbo, Yoruba, Hausa, and lesser-documented tongues. The samples feature rich cultural storytelling, descriptive prose, and the use of reduplication for emphasis. The content ranges from daily life scenes to mythical encounters, serving as a resource for… See the full description on the dataset page: https://huggingface.co/datasets/ChiamakaNwokolo/adaption-african-languages-narratives.textn<1K0 likes16 downloads1mo agoHugging Face19Kamsi08 /african-languages-natlastextn<1K1 likes15 downloads9mo agoHugging Face20African-Languages-Lab /proxy-mt-eval-scoresgated Proxy-MT Eval Scores Corpus-level MT metrics for 50 open-weight LLMs on the translations in proxy-mt-translations. Computed by evaluate_mt.py (BLEU, chrF++, ROUGE-L, METEOR, XCOMET-XL, SSA-COMET). MetricX is backfilled separately and may still be empty in this snapshot. Layout <model>/flores-200.csv <model>/ntrex.csv <model>/wmt24.csv Each CSV has one row per eng-<lang> pair: column description translation-pair e.g. eng-yor bleu sacrebleu corpus BLEU… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-eval-scores.tabulartranslation10K<n<100K0 likes13 downloads1mo agoHugging Face21African-Languages-Lab /all-lab-text-monogatedtext100M<n<1B1 likes12 downloads3mo agoHugging Face22African-Languages-Lab /proxy-mt-benchmark-scoresgated Proxy-MT Benchmark Scores Multilingual benchmark results for 50 open-weight LLMs, evaluated with the lm-evaluation-harness via a vLLM backend. Covers reasoning, comprehension, and knowledge tasks with an emphasis on African and other lower-resource languages. Layout scores/<model>.csv # parsed per-language scores (tidy, ready to plot) raw/<model>/.../results_*.json # raw lm-eval-harness result files raw/<model>/raw_log.txt # full evaluation… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-benchmark-scores.text-generationn<1K0 likes12 downloads1mo agoHugging Face23ouilyh /qwen-blindspots-african-languages Blind Spots of Qwen3.5-0.8B Overview This dataset documents failure cases of the base modelQwen3.5-0.8B. The model has approximately 0.8 billion parameters and is designed as a lightweight multilingual language model. The goal of this dataset is to highlight cases where the model produces incorrect, misleading, or suboptimal outputs across diverse tasks including translation, reasoning, and regional knowledge. The dataset focuses particularly on low-resource African… See the full description on the dataset page: https://huggingface.co/datasets/ouilyh/qwen-blindspots-african-languages.textn<1K0 likes9 downloads6mo agoHugging Face24African-Languages-Lab /all-lab-speechgated All Lab Speech Cleaned African-language speech with embedded, playable audio (HF Audio). One config per language (<lang> = transcribed, <lang>_manifest = audio-only); the audio column sits right after audio_id and plays in the dataset viewer. Splits (train/validation/test) come from the source split labels. from datasets import load_dataset ds = load_dataset("African-Languages-Lab/all-lab-speech", "afrikaans") Columns audio_id, audio (playable), transcript… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/all-lab-speech.audioautomatic-speech-recognition10M<n<100M2 likes4 downloads3mo agoHugging Face25LocaleNLP /afrilion-african-languages0 likes3 downloads5mo agoHugging Face26Uriath /g2p-african-languagesgated G2P African Languages — Mooré · Bambara · Dioula Projet PFE : Leveraging Phonemic Features for Cross-lingual NLP in African LanguagesCITADEL, Ouagadougou, Burkina Faso — 2025-2026 Dataset graphème-to-phonème (G2P) pour trois langues africaines sous-dotées du Burkina Faso. Il couvre le Mooré (famille Gur), le Bambara et le Dioula (famille Mandé), et fournit deux transcriptions phonémiques par entrée : une avec les marques tonales issues du dictionnaire source, une sans tons —… See the full description on the dataset page: https://huggingface.co/datasets/Uriath/g2p-african-languages.text10K<n<100K0 likes3 downloads4mo agoHugging Face27BeardedMonster /african_languages_translation1 likes2 downloads1y agoHugging Face28African-Languages-Lab /all-lab-text-multigatedtabular100M<n<1B1 likes2 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.