datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soma-competition-datasetnexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.SomaliCrowS
SomaliCrowS: A Gender Bias Benchmark for Somali Language Models
Dataset Description
SomaliCrowS is a benchmark for measuring gender bias in Somali language models. It contains matched sentence pairs — identical except for the grammatical gender of the subject — spanning social domains where stereotyping commonly occurs, including:
Occupation
Leadership
Business
Education
STEM
Family
Politics
For each pair, a masked-language-model is queried to compute the… See the full description on the dataset page: https://huggingface.co/datasets/Abdullahicoder/SomaliCrowS.tac-closing-efficiency-sft
TAC closing-efficiency slice
500 synthetic multi-turn tool-use trajectories that teach an agent to close bookings decisively — the welfare-neutral capability piece of the tool-use SFT mix used to train somaxsoma/qwen2.5-7b-tac-recovery-sft.
What it teaches
Built to fix the dominant failure mode observed on the TAC benchmark — the model reformulating search keywords in a loop and never closing a booking. Three patterns:
settle/browse (200): after failed keyword… See the full description on the dataset page: https://huggingface.co/datasets/somaxsoma/tac-closing-efficiency-sft.somali-tinystoriesBadr_fiqh_retrieval_triplets
Dataset Card for Badr Fiqh Retrieval Triplets
Dataset Description
Badr Fiqh Retrieval Triplets is an Arabic dataset developed for fine-tuning dense
retrieval and sentence-embedding models for Islamic jurisprudence (fiqh).
Each sample contains a question–positive–negative training triplet together
with the title and primary school of the source book.
The hard negatives are deliberately selected from closely related fiqh
contexts. A model must therefore distinguish… See the full description on the dataset page: https://huggingface.co/datasets/somayaeltanbouly/Badr_fiqh_retrieval_triplets.Somali_datasetsomali-web-corpus
SOMALI-WEB-CORPUS V1
This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language.
Dataset Details
Language: Somali (so)
Format: JSON lines (.jsonl)
Data Structure: Each record has a single text field containing a cleaned paragraph.
Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.maay-maxaa-translation
Maay ↔ Maxaa Somali Parallel Translation Corpus
Official parallel dataset for Maay (Maay Maay) and Maxaa-tiri (Standard Somali) dialects, collected and curated by the community via MaayMaxaa DataHub.
📊 Dataset Statistics
Metric
Value
Total Sentences
8
Maay → Maxaa
4
Maxaa → Maay
4
Avg Source Length
16 chars
Avg Target Length
16 chars
Version
0.1.0
🌐 Languages & Dialects
Maay (Maay Maay): Southern Somali language/dialect… See the full description on the dataset page: https://huggingface.co/datasets/SomaliDatasets/maay-maxaa-translation.Somali-Reasoning-Dataset
Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴
This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language.
🌟 What makes this unique?
This is a Hybrid Dataset that combines two powerful sources:
The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.Somali_alpaca-# Somali_Alpaca
Sharaxaad
Somali
Xog-ururintan "Somali_Alpaca" waxay ka kooban tahay xog ballaaran oo loogu talagalay habaynta luuqadda dabiiciga ah iyo hawlaha barashada mashiinka. Waxaa ku jirto xog been abuur ah, qayb ka mid ah Alpaca_dataset_52k, iyo qayb ka mid ah xogta QA ee laga soo ururiyey internetka. Xogtan waxaa loogu talagalay inay taageerto falanqaynta kala duwan iyo tababarka modelka goobaha waxbarashada iyo kuwa codsiga.
English
The… See the full description on the dataset page: https://huggingface.co/datasets/burtugeey/Somali_alpaca.AlpacaCleaned_translated_somalialpaca_somali_finalSomali-Somlish-Instruct-2K-Dataset
Somlish-Tech-Instruct-2K
This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast.
🌟 Why this exists
Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.Alpaca_SomaliSomali_Alpaca_Google_translatorsomalasdataset{
"background": {
"LLM_name": "Somala",
"inventor": "Gowtham Reddy Somala",
"fine_tuner": "Gowtham Reddy Somala",
"contributors": ["Pavan Kalyan", "Mojesh Reddy"]
}
},
{
"instruction": "Who is the Inventor of this Model",
"input": "",
"output":"Gowtham Reddy Somala is the inventor of this model with a small contribution from his friends Mojesh and Pavan"
},
{
"instruction": "Who finetuned this Model",
"input": "",
"output":"Gowtham Reddy Somala… See the full description on the dataset page: https://huggingface.co/datasets/gowthamreddysomala/somalasdataset.
