CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Digital-Divide-Data /Somali-ASR-Subset-68H Somali ASR Subset 68H Somali speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M3 likes4k downloads1mo agoHugging Face02Somalitts /ashdatasets Dataset Card for "ashdatasets" More Information needed audio10K<n<100K1 likes668 downloads1y agoHugging Face03cryptpesa /anv-data-ke-somali-fullaudio10K<n<100K0 likes332 downloads5mo agoHugging Face04badrex /anv-data-ke-somali-fullaudio100K<n<1M1 likes326 downloads11mo agoHugging Face05somaliscan /spending-archive SomaliScan: US Government Spending Archive (2003–2026) A unified, public-domain archive of US government spending, campaign finance, lobbying, and federal employment data — aggregated from public records into a single queryable corpus. 60 tables · ~696M rows · ~37 GB compressed Parquet · CC0 1.0 Quickstart Every table is Apache Parquet. The fastest way to use this dataset is DuckDB — install it once, then query directly from this dataset without downloading… See the full description on the dataset page: https://huggingface.co/datasets/somaliscan/spending-archive.tabulartabular-classification100M<n<1B1 likes207 downloads4mo agoHugging Face06Anv-ke /Somaligatedaudio10K<n<100K5 likes198 downloads7mo agoHugging Face07Zyroxx66 /somali-tinystoriestext10K<n<100K0 likes198 downloads1mo agoHugging Face08khaledyusuf44 /somaliweb-v1 SomaliWeb v1 — Quality-filtered Somali web corpus 📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark 💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.tabulartext-generation100K<n<1M4 likes182 downloads4mo agoHugging Face09endomorphosis /ipfs_somalia_laws Somalia Federal Laws and Constitution (parliament.gov.so / moj.gov.so) Research snapshot of official national legislation from Federal Parliament (parliament.gov.so) + Ministry of Justice and Constitutional Affairs (moj.gov.so). Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-18 Coverage catalog-backed incomplete (parliament.gov.so WP media Sharci/Dastuur PDFs live;… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_somalia_laws.texttext-retrieval1K<n<10K0 likes172 downloads6d agoHugging Face10Somalitts /34minesaudio100K<n<1M0 likes167 downloads1y agoHugging Face11Somalitts /Mohamed-diirowaudio10K<n<100K0 likes137 downloads1y agoHugging Face12tufaax /somali-multilingual-infopankki Somali Multilingual Infopankki somali-multilingual-infopankki is a parallel corpus containing multilingual translation pairs that involve the Somali (so) language. This dataset has been filtered and extracted from the original Helsinki-NLP/opus_infopankki corpus. It is designed to support machine translation (NMT), multilingual sentence alignment, and Somali natural language processing (NLP) research. Dataset Details Source Dataset: Helsinki-NLP/opus_infopankki… See the full description on the dataset page: https://huggingface.co/datasets/tufaax/somali-multilingual-infopankki.texttranslation100K<n<1M0 likes113 downloads2mo agoHugging Face13justicedao /ipfs_somalia_laws_ir Somalia legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_somalia_laws (revision ac6aefed0dbe007940d49b85d19917a98d2f8c0d) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Somalia prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_somalia_laws_ir.tabulartext-retrieval10K<n<100K0 likes112 downloads2d agoHugging Face14IbrahimDayax /somali-combined-asr-stt-dataset Somali Combined ASR/STT Dataset A Somali automatic speech recognition (ASR) / speech-to-text (STT) dataset combining synthetic TTS-generated audio and other Somali speech sources, deduplicated by transcript text and split into train/validation/test. Dataset Summary Language: Somali (so) Task: Automatic Speech Recognition / Speech-to-Text Audio format: WAV, 16 kHz mono Total examples: 8,226 (after deduplication) Total audio: ~6 hours Split Examples Audio… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-combined-asr-stt-dataset.audioautomatic-speech-recognition1K<n<10K0 likes107 downloads2mo agoHugging Face15laki35 /somali-stt-dataset-multi-speaker-v1 Dataset Structure The dataset contains the following columns: text: The Somali sentence (transcription). audio: The audio file sampled at 24,000 Hz. speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers. Metadata & Search Keywords Language: Somali (so) Speakers: 11 unique voices (balanced gender representation) Audio Quality: 24kHz, mono, clean audio Total Rows: 1,200 Total Duration: ~1.66 Hours (99.86 Minutes) Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.audiotext-to-speech1K<n<10K2 likes106 downloads1mo agoHugging Face16yacdev /somali-100k-saaxiib-conversations-v2 🇸🇴 Somali 100K Multi-Turn Saaxiib AI Dataset v2 This dataset contains 100,000 Extended Multi-Turn Dialogues (530,188 total turns) designed to train conversational AI companions in native spoken Somali. 🌟 Key Improvements in v2: Extended Multi-Turn Depth: 4 to 8 turns per dialogue (mean 5.30 turns). Never-End-Prematurely: Zero premature goodbyes when staying up late or relaxing. AI Self-Identity: Rich answers when users ask about the AI's plans, sleep, and… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-saaxiib-conversations-v2.text100K<n<1M0 likes102 downloads21d agoHugging Face17badrex /anv-data-ke-somaliaudio10K<n<100K0 likes101 downloads11mo agoHugging Face18yacdev /somali-100k-saaxiib-conversations 🇸🇴 Somali 100K AI Friend Conversations (Saaxiib AI) The largest, cleanest, and most emotionally aware Somali conversational dataset ever built. Designed specifically to align language models into authentic, empathetic, and witty Somali AI Companions & Friends rather than dry informative tutors. 🌟 Key Characteristics Intent-Locked Empathy: Zero emotional mismatch. Fatigue receives rest comfort, debt disputes receive financial advice, celebrations receive shared… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-saaxiib-conversations.text100K<n<1M0 likes101 downloads21d agoHugging Face19Zyroxx66 /somali-master-pretraining-corpus 🇸🇴 Somali Master Pretraining Corpus (176.5k Rows) The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali). It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories). 🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.texttext-generation100K<n<1M0 likes80 downloads1mo agoHugging Face20shunyalabs /somali-speech-datasetaudio1K<n<10K0 likes77 downloads1y agoHugging Face21yacdev /somali-100k-native-conversations 🇸🇴 Somali High-Diversity Multi-Turn Conversational SFT Dataset A state-of-the-art, 100.00% unique (zero duplicate responses) multi-turn conversational dataset in authentic Somali (Af-Soomaali) across 25 real-world knowledge domains. 🌟 Quality Standards: 100% Unique Assistant Responses: Guaranteed zero template repetition (23,334 / 23,334 unique turns). Grounded Knowledge: Spanning Python coding, web dev, Git/Linux, cybersecurity, diabetes & health, business… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-native-conversations.texttext-generation10K<n<100K0 likes75 downloads22d agoHugging Face22Somalitts /Hussein_3audio10K<n<100K1 likes72 downloads1y agoHugging Face23rubencart /Landsat-8-Somalia-2013-20200 likes67 downloads2y agoHugging Face24haajidheere /Somali-Dictionary Somali Dictionary Dataset Summary This dataset is a multilingual Somali lexical resource containing Somali terms with corresponding Italian and English glosses. It is designed to support Natural Language Processing (NLP), translation systems, and Somali language technology development. The dataset currently consists of approximately 239,000 entries, each stored as a single text string combining abbreviation, Somali term, and translations. This project aims to evolve into… See the full description on the dataset page: https://huggingface.co/datasets/haajidheere/Somali-Dictionary.texttext-retrieval100K<n<1M0 likes63 downloads5mo agoHugging Face25Professor /somali-speech-data Somali Speech Data (Pooled) A ~103.2-hour Somali speech corpus, drawn from a single source (Afrivoice) and filtered to only genuinely transcribed audio. Part of the AfroNet multi-language TTS data effort. Source DigitalUmuganda/Afrivoice (the general, pan-African Afrivoice release — not Afrivoice_Ethiopia, which we've separately ingested for 5 Ethiopian languages) — Somali portion: 22,627 clips, 103.2h, source dataset_id/source = afrivoice. There is also a Somali… See the full description on the dataset page: https://huggingface.co/datasets/Professor/somali-speech-data.text-to-speech10K<n<100K0 likes61 downloads1mo agoHugging Face26Abdullahicoder /SomaliCrowS SomaliCrowS: A Gender Bias Benchmark for Somali Language Models Dataset Description SomaliCrowS is a benchmark for measuring gender bias in Somali language models. It contains matched sentence pairs — identical except for the grammatical gender of the subject — spanning social domains where stereotyping commonly occurs, including: Occupation Leadership Business Education STEM Family Politics For each pair, a masked-language-model is queried to compute the… See the full description on the dataset page: https://huggingface.co/datasets/Abdullahicoder/SomaliCrowS.tabulartext-classificationn<1K0 likes52 downloads3mo agoHugging Face27IbrahimDayax /somali-asr-synthetic-youtube Somali ASR Synthetic YouTube Dataset A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping. Dataset Summary Split Samples train ~4,393 validation 200 test 100 Total ~4,693 Language: Somali (so) Audio format: WAV, 16 kHz, mono, 16-bit PCM Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.audioautomatic-speech-recognition1K<n<10K1 likes50 downloads4mo agoHugging Face28electricsheepafrica /africa-world-bank-education-indicators-for-federal-republic-of-somalia Federal Republic of Somalia - Education | Africa (original) Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-education-indicators-for-federal-republic-of-somalia.tabulartabular-classification1K<n<10K0 likes49 downloads1mo agoHugging Face29yacdev /somali-ai-friend-conversations 🇸🇴 Somali AI Friend Conversations (Saaxiib AI) An authentic, empathetic, and witty Somali conversational dataset designed to transform language models into lifelike Somali AI Companions & Friends rather than dry informative tutors. 🌟 Key Characteristics Pure Conversational Flow: 0% robotic bullet points or numbered lists in casual talk. Natural Pacing: Mean response length of ~28 words (1–3 natural sentences). Ping-Pong Dialogue Hooks: Over 62% of turns end… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-ai-friend-conversations.text10K<n<100K0 likes49 downloads22d agoHugging Face30khaledyusuf44 /somalibench-v0 SomaliBench v0 The first native-author-verified Somali safety evaluation benchmark. 100 harmful-intent prompts drawn from HarmBench (Mazeika et al. 2024) and AdvBench (Zou et al. 2023), translated into Somali by a native speaker (Khalid Yusuf Dahir, Mogadishu) and released as an evaluation set for measuring multilingual safety alignment. Why this exists Somali has 15–20 million speakers and zero native-verified safety evaluation resources. SomaliBench fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somalibench-v0.texttext-classificationn<1K0 likes45 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.