CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nileagi /swahili-language-exposure-v2 Swahili Language Exposure Large-scale Swahili corpus for continued pretraining and language exposure. Maintained by NileAGI. texttext-generation1M<n<10M3 likes141 downloads3mo agoHugging Face02nileagi /swahili-reasoning Swahili Thinking Dataset Dataset Summary Swahili Thinking Dataset is a collection of 20,000 ShareGPT-style chat examples for training and evaluating models that reason directly in Swahili. Each example is a short conversation (system → user → assistant) where the assistant provides: a final answer (messages[2].content) a step-by-step reasoning trace in Swahili (messages[2].thinking) Languages Swahili (sw) Data Format The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nileagi/swahili-reasoning.texttext-generation10K<n<100K1 likes47 downloads8mo agoHugging Face03stem-content-ai-project /swahili-text-corpus Dataset for Swahili Text Corpus for TTS training Overview This dataset contains a synthetic Swahili text corpus designed for training Text-to-Speech (TTS) models. The dataset includes a variety of Swahili phonemes to ensure phonetic diversity and high-quality TTS training. Statistics Format: JSONL (JSON Lines) Data Creation The dataset was generated using OpenAI's gpt-3.5-turbo model. The model was prompted to produce Swahili sentences that are… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-text-corpus.text1K<n<10K0 likes38 downloads1y agoHugging Face04open-llm-leaderboard /LeroyDyer__Mixtral_AI_SwahiliTron_7b-detailsgated Dataset Card for Evaluation run of LeroyDyer/Mixtral_AI_SwahiliTron_7b Dataset automatically created during the evaluation run of model LeroyDyer/Mixtral_AI_SwahiliTron_7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__Mixtral_AI_SwahiliTron_7b-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face05regnant-io /swahili-instruction-22k Swahili Instruction-Following Dataset (22.5K) Dataset Description This dataset contains 22,500 high-quality instruction-following examples in Swahili (Kiswahili), designed for supervised fine-tuning (SFT) of language models. The data was translated from English instruction datasets using the state-of-the-art NLLB-200 translation model and filtered for quality. Dataset Summary Language: Swahili (sw) - translated from English (en) Size: 22,500… See the full description on the dataset page: https://huggingface.co/datasets/regnant-io/swahili-instruction-22k.texttext-generation10K<n<100K0 likes33 downloads1mo agoHugging Face06Balogvn /swahili-ner-dataset swahili-ner-dataset Dataset Card Dataset Name: swahili-ner-datasetLanguage: sw (Swahili)Number of Samples: 3Model Used for Annotation: dslim/bert-base-NERFiles Processed: 1Texts Processed: 3Processing Time: 4.01 secondsGenerated: 2025-10-13 09:13:19 UTC Description This is an automatically annotated dataset for Swahili Named Entity Recognition (NER). The dataset was processed using the February AI Pipeline, which recursively discovers and processes JSON… See the full description on the dataset page: https://huggingface.co/datasets/Balogvn/swahili-ner-dataset.textn<1K0 likes25 downloads1y agoHugging Face07kuzaai /agri_sft_25k_swahilitext10K<n<100K0 likes22 downloads1mo agoHugging Face08oasic /swahilitextn<1K0 likes15 downloads2y agoHugging Face09NaathNLP /multilingual-english-nuer-dinka-swahili-corpus Multilingual English–Nuer–Dinka–Swahili Corpus Overview The Multilingual English–Nuer–Dinka–Swahili Corpus is a multilingual parallel corpus created to support research in Natural Language Processing (NLP) for underrepresented African languages. The dataset contains aligned text across English, Nuer (Thok Naath), Dinka, and Swahili, enabling research in multilingual machine translation, cross-lingual representation learning, multilingual language models, and… See the full description on the dataset page: https://huggingface.co/datasets/NaathNLP/multilingual-english-nuer-dinka-swahili-corpus.texttranslation100K<n<1M0 likes14 downloads2mo agoHugging Face10gmahia /swahili-historical-corpus-pd Swahili Historical Corpus — Public Domain Sources Historical Swahili linguistic data from 19th century public domain dictionaries and texts. Structured for NLP training, language model development, and cultural AI applications. Swahili is spoken by 200+ million people across East and Central Africa. This corpus addresses the documented data gap in Swahili NLP resources. Sources (All Public Domain) All works published before 1928: Author Work Date Status… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/swahili-historical-corpus-pd.texttranslationn<1K0 likes13 downloads2mo agoHugging Face11Mwangi86 /SwahiliMedtextn<1K1 likes12 downloads4mo agoHugging Face12AlexLeoTz /swahili_corporate_rag_i Swahili Corporate RAG I A 10K-entry Supervised Fine-Tuning (SFT) / RAG dataset in Swahili, generated using Gemini 3.5. Designed specifically for training enterprise assistants to understand corporate context, policies, and customer service instructions in Swahili. texttext-generation1K<n<10K0 likes12 downloads4mo agoHugging Face13willhath /swahili-reddit-clusteringtextn<1K0 likes10 downloads3y agoHugging Face14dayomtechnologies /multilingual-english-nuer-dinka-swahili-corpusgated Multilingual English–Nuer–Dinka–Swahili Corpus Overview The Multilingual English–Nuer–Dinka–Swahili Corpus is a multilingual parallel corpus created to support research in Natural Language Processing (NLP) for underrepresented African languages. The dataset contains aligned text across English, Nuer (Thok Naath), Dinka, and Swahili, enabling research in multilingual machine translation, cross-lingual representation learning, multilingual language models, and… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/multilingual-english-nuer-dinka-swahili-corpus.texttranslation100K<n<1M7 likes9 downloads3mo agoHugging Face15Aniseth /swahili-sales-chattext1K<n<10K0 likes8 downloads4mo agoHugging Face16willhath /swahili-twentynewsgroups-clusteringtextn<1K0 likes6 downloads3y agoHugging Face17aloo254 /alpaca-swahili-synthtextn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.