CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01khaledyusuf44 /somaliweb-v1 SomaliWeb v1 — Quality-filtered Somali web corpus 📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark 💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.tabulartext-generation100K<n<1M4 likes170 downloads4mo agoHugging Face02yacdev /somali-100k-native-conversations 🇸🇴 Somali High-Diversity Multi-Turn Conversational SFT Dataset A state-of-the-art, 100.00% unique (zero duplicate responses) multi-turn conversational dataset in authentic Somali (Af-Soomaali) across 25 real-world knowledge domains. 🌟 Quality Standards: 100% Unique Assistant Responses: Guaranteed zero template repetition (23,334 / 23,334 unique turns). Grounded Knowledge: Spanning Python coding, web dev, Git/Linux, cybersecurity, diabetes & health, business… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-native-conversations.texttext-generation10K<n<100K0 likes77 downloads24d agoHugging Face03Zyroxx66 /somali-master-pretraining-corpus 🇸🇴 Somali Master Pretraining Corpus (176.5k Rows) The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali). It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories). 🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.texttext-generation100K<n<1M0 likes42 downloads1mo agoHugging Face04khaledyusuf44 /somalibench-v0 SomaliBench v0 The first native-author-verified Somali safety evaluation benchmark. 100 harmful-intent prompts drawn from HarmBench (Mazeika et al. 2024) and AdvBench (Zou et al. 2023), translated into Somali by a native speaker (Khalid Yusuf Dahir, Mogadishu) and released as an evaluation set for measuring multilingual safety alignment. Why this exists Somali has 15–20 million speakers and zero native-verified safety evaluation resources. SomaliBench fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somalibench-v0.texttext-classificationn<1K0 likes36 downloads4mo agoHugging Face05michsethowusu /Code-170k-somali Dataset Description Code-170k-somali is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Somali, making coding education accessible to Somali speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Somali language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-somali.texttext-generation100K<n<1M1 likes32 downloads11mo agoHugging Face06IbraahimLab /fineweb-somali FineWeb Somali Dataset A curated collection of Somali language content from BBC Somali, designed for training and evaluating small language models on low-resource languages. Dataset Description This dataset contains 4,910 high-quality Somali language articles scraped from BBC Somali's website. The content covers diverse topics including news, culture, technology, sports, and human interest stories, providing a rich corpus for Somali language model training.… See the full description on the dataset page: https://huggingface.co/datasets/IbraahimLab/fineweb-somali.texttext-generation1K<n<10K0 likes28 downloads8mo agoHugging Face07Zyroxx66 /Somali-Reasoning-Dataset Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴 This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language. 🌟 What makes this unique? This is a Hybrid Dataset that combines two powerful sources: The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.texttext-generation10K<n<100K0 likes22 downloads6mo agoHugging Face08maanka2 /somali-web-corpus SOMALI-WEB-CORPUS V1 This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language. Dataset Details Language: Somali (so) Format: JSON lines (.jsonl) Data Structure: Each record has a single text field containing a cleaned paragraph. Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.texttext-generation100K<n<1M1 likes20 downloads4mo agoHugging Face09Zyroxx66 /somali-pretraining-corpus 🇸🇴 Somali Open Pre-Training Corpus v2 (SOPC-v2) Creator: Hamze Jamal (@Zyroxx66)Language: Somali (Af-Soomaali)License: CC-BY-4.0Release Version: 2.0 📌 Overview The Somali Open Pre-Training Corpus v2 (SOPC-v2) is a massively expanded, high-quality, deduplicated, and domain-balanced dataset designed specifically for Continued Pre-Training (CPT), Domain Adaptation, and Instruction Alignment of Large Language Models (LLMs) such as Qwen 2.5, Llama 3.2, Gemma 2… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-pretraining-corpus.texttext-generation10K<n<100K0 likes16 downloads2mo agoHugging Face10Zyroxx66 /Somali-Somlish-Instruct-2K-Dataset Somlish-Tech-Instruct-2K This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast. 🌟 Why this exists Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.texttext-generation1K<n<10K0 likes13 downloads6mo agoHugging Face11jojo-ai-mst /Roleplay-Somali RolePlay-Somali Roleplay-Somali Dataset is a dataset for roleplaying in the Somali language for Large Language Model. The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, see this github repo. For contact… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Somali.texttext-generation1K<n<10K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.