datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nexttoken-model-1-dataset-sft
NextToken Model 1 SFT dataset (v4)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
846 scheme/product source documents across 57 schemes/products, chunked
into 1,445 passages.
v4 vs v3: v3 merged in a second batch (6,664 rows) without… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-1-dataset-sft.soma-skincare-qanexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.SOMAJGYAAN
SomajGyaan (সমাজজ্ঞান) - Bangla MCQ Dataset
📊 Dataset Description
SomajGyaan (সমাজজ্ঞান) is a comprehensive Bangla multiple-choice question dataset featuring 4,234 questions across 7 academic categories with ~12,000 unique answer options.
Dataset Summary
Total Questions: 4,234
Unique Answer Options: ~12,000
Answer Diversity: 70.8%
Language: Bangla (Bengali)
Categories: 7 (History, Economics, Geography, Politics, Social Studies, Law… See the full description on the dataset page: https://huggingface.co/datasets/farihashifa/SOMAJGYAAN.Code-170k-somali
Dataset Description
Code-170k-somali is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Somali, making coding education accessible to Somali speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Somali language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-somali.Somali-Reasoning-Dataset
Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴
This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language.
🌟 What makes this unique?
This is a Hybrid Dataset that combines two powerful sources:
The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.Somali-Somlish-Instruct-2K-Dataset
Somlish-Tech-Instruct-2K
This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast.
🌟 Why this exists
Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.
