datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
ashdatasets
Dataset Card for "ashdatasets"
More Information needed
anv-data-ke-somali-fullanv-data-ke-somali-fullspending-archive
SomaliScan: US Government Spending Archive (2003–2026)
A unified, public-domain archive of US government spending, campaign
finance, lobbying, and federal employment data — aggregated from public
records into a single queryable corpus.
60 tables · ~696M rows · ~37 GB compressed Parquet · CC0 1.0
Quickstart
Every table is Apache Parquet. The fastest way to use this dataset is
DuckDB — install it once, then query directly
from this dataset without downloading… See the full description on the dataset page: https://huggingface.co/datasets/somaliscan/spending-archive.Somalisomali-tinystoriessomaliweb-v1
SomaliWeb v1 — Quality-filtered Somali web corpus
📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus
SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.ipfs_somalia_laws
Somalia Federal Laws and Constitution (parliament.gov.so / moj.gov.so)
Research snapshot of official national legislation from Federal Parliament (parliament.gov.so) + Ministry of Justice and Constitutional Affairs (moj.gov.so).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-18
Coverage
catalog-backed incomplete (parliament.gov.so WP media Sharci/Dastuur PDFs live;… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_somalia_laws.34minesMohamed-diirowsomali-multilingual-infopankki
Somali Multilingual Infopankki
somali-multilingual-infopankki is a parallel corpus containing multilingual translation pairs that involve the Somali (so) language. This dataset has been filtered and extracted from the original Helsinki-NLP/opus_infopankki corpus.
It is designed to support machine translation (NMT), multilingual sentence alignment, and Somali natural language processing (NLP) research.
Dataset Details
Source Dataset: Helsinki-NLP/opus_infopankki… See the full description on the dataset page: https://huggingface.co/datasets/tufaax/somali-multilingual-infopankki.ipfs_somalia_laws_ir
Somalia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_somalia_laws (revision ac6aefed0dbe007940d49b85d19917a98d2f8c0d) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Somalia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_somalia_laws_ir.somali-combined-asr-stt-dataset
Somali Combined ASR/STT Dataset
A Somali automatic speech recognition (ASR) / speech-to-text (STT) dataset combining
synthetic TTS-generated audio and other Somali speech sources, deduplicated by transcript
text and split into train/validation/test.
Dataset Summary
Language: Somali (so)
Task: Automatic Speech Recognition / Speech-to-Text
Audio format: WAV, 16 kHz mono
Total examples: 8,226 (after deduplication)
Total audio: ~6 hours
Split
Examples
Audio… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-combined-asr-stt-dataset.somali-stt-dataset-multi-speaker-v1
Dataset Structure
The dataset contains the following columns:
text: The Somali sentence (transcription).
audio: The audio file sampled at 24,000 Hz.
speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers.
Metadata & Search Keywords
Language: Somali (so)
Speakers: 11 unique voices (balanced gender representation)
Audio Quality: 24kHz, mono, clean audio
Total Rows: 1,200
Total Duration: ~1.66 Hours (99.86 Minutes)
Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.somali-100k-saaxiib-conversations-v2
🇸🇴 Somali 100K Multi-Turn Saaxiib AI Dataset v2
This dataset contains 100,000 Extended Multi-Turn Dialogues (530,188 total turns) designed to train conversational AI companions in native spoken Somali.
🌟 Key Improvements in v2:
Extended Multi-Turn Depth: 4 to 8 turns per dialogue (mean 5.30 turns).
Never-End-Prematurely: Zero premature goodbyes when staying up late or relaxing.
AI Self-Identity: Rich answers when users ask about the AI's plans, sleep, and… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-saaxiib-conversations-v2.anv-data-ke-somalisomali-100k-saaxiib-conversations
🇸🇴 Somali 100K AI Friend Conversations (Saaxiib AI)
The largest, cleanest, and most emotionally aware Somali conversational dataset ever built. Designed specifically to align language models into authentic, empathetic, and witty Somali AI Companions & Friends rather than dry informative tutors.
🌟 Key Characteristics
Intent-Locked Empathy: Zero emotional mismatch. Fatigue receives rest comfort, debt disputes receive financial advice, celebrations receive shared… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-saaxiib-conversations.somali-master-pretraining-corpus
🇸🇴 Somali Master Pretraining Corpus (176.5k Rows)
The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali).
It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories).
🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.somali-speech-datasetsomali-100k-native-conversations
🇸🇴 Somali High-Diversity Multi-Turn Conversational SFT Dataset
A state-of-the-art, 100.00% unique (zero duplicate responses) multi-turn conversational dataset in authentic Somali (Af-Soomaali) across 25 real-world knowledge domains.
🌟 Quality Standards:
100% Unique Assistant Responses: Guaranteed zero template repetition (23,334 / 23,334 unique turns).
Grounded Knowledge: Spanning Python coding, web dev, Git/Linux, cybersecurity, diabetes & health, business… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-native-conversations.Hussein_3Landsat-8-Somalia-2013-2020Somali-Dictionary
Somali Dictionary
Dataset Summary
This dataset is a multilingual Somali lexical resource containing Somali terms with corresponding Italian and English glosses. It is designed to support Natural Language Processing (NLP), translation systems, and Somali language technology development.
The dataset currently consists of approximately 239,000 entries, each stored as a single text string combining abbreviation, Somali term, and translations.
This project aims to evolve into… See the full description on the dataset page: https://huggingface.co/datasets/haajidheere/Somali-Dictionary.somali-speech-data
Somali Speech Data (Pooled)
A ~103.2-hour Somali speech corpus, drawn from a single source (Afrivoice) and
filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort.
Source
DigitalUmuganda/Afrivoice
(the general, pan-African Afrivoice release — not Afrivoice_Ethiopia, which we've
separately ingested for 5 Ethiopian languages) — Somali portion: 22,627 clips,
103.2h, source dataset_id/source = afrivoice.
There is also a Somali… See the full description on the dataset page: https://huggingface.co/datasets/Professor/somali-speech-data.SomaliCrowS
SomaliCrowS: A Gender Bias Benchmark for Somali Language Models
Dataset Description
SomaliCrowS is a benchmark for measuring gender bias in Somali language models. It contains matched sentence pairs — identical except for the grammatical gender of the subject — spanning social domains where stereotyping commonly occurs, including:
Occupation
Leadership
Business
Education
STEM
Family
Politics
For each pair, a masked-language-model is queried to compute the… See the full description on the dataset page: https://huggingface.co/datasets/Abdullahicoder/SomaliCrowS.somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.africa-world-bank-education-indicators-for-federal-republic-of-somalia
Federal Republic of Somalia - Education | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-education-indicators-for-federal-republic-of-somalia.somali-ai-friend-conversations
🇸🇴 Somali AI Friend Conversations (Saaxiib AI)
An authentic, empathetic, and witty Somali conversational dataset designed to transform language models into lifelike Somali AI Companions & Friends rather than dry informative tutors.
🌟 Key Characteristics
Pure Conversational Flow: 0% robotic bullet points or numbered lists in casual talk.
Natural Pacing: Mean response length of ~28 words (1–3 natural sentences).
Ping-Pong Dialogue Hooks: Over 62% of turns end… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-ai-friend-conversations.somalibench-v0
SomaliBench v0
The first native-author-verified Somali safety evaluation benchmark.
100 harmful-intent prompts drawn from HarmBench (Mazeika et al. 2024) and
AdvBench (Zou et al. 2023), translated into Somali by a native speaker
(Khalid Yusuf Dahir, Mogadishu) and released as an evaluation set for
measuring multilingual safety alignment.
Why this exists
Somali has 15–20 million speakers and zero native-verified safety
evaluation resources. SomaliBench fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somalibench-v0.
