datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
somaliweb-v1
SomaliWeb v1 — Quality-filtered Somali web corpus
📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus
SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.somali-100k-native-conversations
🇸🇴 Somali High-Diversity Multi-Turn Conversational SFT Dataset
A state-of-the-art, 100.00% unique (zero duplicate responses) multi-turn conversational dataset in authentic Somali (Af-Soomaali) across 25 real-world knowledge domains.
🌟 Quality Standards:
100% Unique Assistant Responses: Guaranteed zero template repetition (23,334 / 23,334 unique turns).
Grounded Knowledge: Spanning Python coding, web dev, Git/Linux, cybersecurity, diabetes & health, business… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-native-conversations.somali-master-pretraining-corpus
🇸🇴 Somali Master Pretraining Corpus (176.5k Rows)
The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali).
It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories).
🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.somalibench-v0
SomaliBench v0
The first native-author-verified Somali safety evaluation benchmark.
100 harmful-intent prompts drawn from HarmBench (Mazeika et al. 2024) and
AdvBench (Zou et al. 2023), translated into Somali by a native speaker
(Khalid Yusuf Dahir, Mogadishu) and released as an evaluation set for
measuring multilingual safety alignment.
Why this exists
Somali has 15–20 million speakers and zero native-verified safety
evaluation resources. SomaliBench fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somalibench-v0.Code-170k-somali
Dataset Description
Code-170k-somali is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Somali, making coding education accessible to Somali speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Somali language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-somali.fineweb-somali
FineWeb Somali Dataset
A curated collection of Somali language content from BBC Somali, designed for training and evaluating small language models on low-resource languages.
Dataset Description
This dataset contains 4,910 high-quality Somali language articles scraped from BBC Somali's website. The content covers diverse topics including news, culture, technology, sports, and human interest stories, providing a rich corpus for Somali language model training.… See the full description on the dataset page: https://huggingface.co/datasets/IbraahimLab/fineweb-somali.Somali-Reasoning-Dataset
Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴
This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language.
🌟 What makes this unique?
This is a Hybrid Dataset that combines two powerful sources:
The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.somali-web-corpus
SOMALI-WEB-CORPUS V1
This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language.
Dataset Details
Language: Somali (so)
Format: JSON lines (.jsonl)
Data Structure: Each record has a single text field containing a cleaned paragraph.
Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.somali-pretraining-corpus
🇸🇴 Somali Open Pre-Training Corpus v2 (SOPC-v2)
Creator: Hamze Jamal (@Zyroxx66)Language: Somali (Af-Soomaali)License: CC-BY-4.0Release Version: 2.0
📌 Overview
The Somali Open Pre-Training Corpus v2 (SOPC-v2) is a massively expanded, high-quality, deduplicated, and domain-balanced dataset designed specifically for Continued Pre-Training (CPT), Domain Adaptation, and Instruction Alignment of Large Language Models (LLMs) such as Qwen 2.5, Llama 3.2, Gemma 2… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-pretraining-corpus.Somali-Somlish-Instruct-2K-Dataset
Somlish-Tech-Instruct-2K
This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast.
🌟 Why this exists
Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.Roleplay-Somali
RolePlay-Somali
Roleplay-Somali Dataset is a dataset for roleplaying in the Somali language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, see this github repo.
For contact… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Somali.
