datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swahili-language-exposure-v2
Swahili Language Exposure
Large-scale Swahili corpus for continued pretraining and language exposure.
Maintained by NileAGI.
swahili-reasoning
Swahili Thinking Dataset
Dataset Summary
Swahili Thinking Dataset is a collection of 20,000 ShareGPT-style chat examples for training and evaluating models that reason directly in Swahili.
Each example is a short conversation (system → user → assistant) where the assistant provides:
a final answer (messages[2].content)
a step-by-step reasoning trace in Swahili (messages[2].thinking)
Languages
Swahili (sw)
Data Format
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nileagi/swahili-reasoning.swahili-text-corpus
Dataset for Swahili Text Corpus for TTS training
Overview
This dataset contains a synthetic Swahili text corpus designed for training Text-to-Speech (TTS) models. The dataset includes a variety of Swahili phonemes to ensure phonetic diversity and high-quality TTS training.
Statistics
Format: JSONL (JSON Lines)
Data Creation
The dataset was generated using OpenAI's gpt-3.5-turbo model. The model was prompted to produce Swahili sentences that are… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-text-corpus.LeroyDyer__Mixtral_AI_SwahiliTron_7b-details
Dataset Card for Evaluation run of LeroyDyer/Mixtral_AI_SwahiliTron_7b
Dataset automatically created during the evaluation run of model LeroyDyer/Mixtral_AI_SwahiliTron_7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__Mixtral_AI_SwahiliTron_7b-details.swahili-instruction-22k
Swahili Instruction-Following Dataset (22.5K)
Dataset Description
This dataset contains 22,500 high-quality instruction-following examples in Swahili (Kiswahili), designed for supervised fine-tuning (SFT) of language models. The data was translated from English instruction datasets using the state-of-the-art NLLB-200 translation model and filtered for quality.
Dataset Summary
Language: Swahili (sw) - translated from English (en)
Size: 22,500… See the full description on the dataset page: https://huggingface.co/datasets/regnant-io/swahili-instruction-22k.swahili-ner-dataset
swahili-ner-dataset
Dataset Card
Dataset Name: swahili-ner-datasetLanguage: sw (Swahili)Number of Samples: 3Model Used for Annotation: dslim/bert-base-NERFiles Processed: 1Texts Processed: 3Processing Time: 4.01 secondsGenerated: 2025-10-13 09:13:19 UTC
Description
This is an automatically annotated dataset for Swahili Named Entity Recognition (NER). The dataset was processed using the February AI Pipeline, which recursively discovers and processes JSON… See the full description on the dataset page: https://huggingface.co/datasets/Balogvn/swahili-ner-dataset.agri_sft_25k_swahiliswahilimultilingual-english-nuer-dinka-swahili-corpus
Multilingual English–Nuer–Dinka–Swahili Corpus
Overview
The Multilingual English–Nuer–Dinka–Swahili Corpus is a multilingual parallel corpus created to support research in Natural Language Processing (NLP) for underrepresented African languages.
The dataset contains aligned text across English, Nuer (Thok Naath), Dinka, and Swahili, enabling research in multilingual machine translation, cross-lingual representation learning, multilingual language models, and… See the full description on the dataset page: https://huggingface.co/datasets/NaathNLP/multilingual-english-nuer-dinka-swahili-corpus.swahili-historical-corpus-pd
Swahili Historical Corpus — Public Domain Sources
Historical Swahili linguistic data from 19th century public domain dictionaries and texts.
Structured for NLP training, language model development, and cultural AI applications.
Swahili is spoken by 200+ million people across East and Central Africa.
This corpus addresses the documented data gap in Swahili NLP resources.
Sources (All Public Domain)
All works published before 1928:
Author
Work
Date
Status… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/swahili-historical-corpus-pd.SwahiliMedswahili_corporate_rag_i
Swahili Corporate RAG I
A 10K-entry Supervised Fine-Tuning (SFT) / RAG dataset in Swahili, generated using Gemini 3.5.
Designed specifically for training enterprise assistants to understand corporate context, policies, and customer service instructions in Swahili.
swahili-reddit-clusteringmultilingual-english-nuer-dinka-swahili-corpus
Multilingual English–Nuer–Dinka–Swahili Corpus
Overview
The Multilingual English–Nuer–Dinka–Swahili Corpus is a multilingual parallel corpus created to support research in Natural Language Processing (NLP) for underrepresented African languages.
The dataset contains aligned text across English, Nuer (Thok Naath), Dinka, and Swahili, enabling research in multilingual machine translation, cross-lingual representation learning, multilingual language models, and… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/multilingual-english-nuer-dinka-swahili-corpus.swahili-sales-chatswahili-twentynewsgroups-clusteringalpaca-swahili-synth
