datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smart-contract-aggregators-educationalInstructGpt-educational
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-educational.hermis_singleTurnNepali_educationaleducational_vedantu_FAQs_Nepali_sft_dataset
Vedantu CUET FAQs — Nepali SFT Dataset
A single-turn instruction-following (Q&A) dataset in Nepali, built from Vedantu's CUET (Common University Entrance Test) 2026 FAQ page. The dataset is formatted for supervised fine-tuning (SFT) in a conversations-style chat schema.
Summary
Records
49
Language
Nepali (ne / npi)
Script
Devanagari (Deva)
Format
JSONL, conversations (human/human turns)
Domain
Education — CUET exam FAQs
Task type… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_vedantu_FAQs_Nepali_sft_dataset.educational_domain_dataset
Nepali Grounded Education QA (OpenHermes-format)
A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about
student enrollment statistics from Nepal's Ministry of Education. Every answer is
anchored to a real numeric value pulled from government open data — nothing in the
answers is model-hallucinated.
Dataset Summary
Rows
611
Language
Nepali (Devanagari script)
Format
ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.Annex_7_Educational_Indicators_2024-2025
Grounded Educational Indicators — Nepali SFT Dataset
File: grounded_education_indicators_nepali_sft.jsonl
Records: 10,091
Language: Nepali (ne / ISO 639-3 npi), Devanagari script
License: Apache-2.0 (permissive)
Format: JSON Lines, ShareGPT-style conversational (human / gpt turns)
Task type: Grounded question answering over structured (tabular) educational statistics
Size on disk: ~15.5 MB
1. Overview
This dataset is a supervised fine-tuning (SFT) corpus of 10… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Annex_7_Educational_Indicators_2024-2025.Educational-Lecture-Dataseteducational-math-instruction-dataset
Educational Math Instruction Dataset
Overview
This dataset contains instruction-style examples designed for fine-tuning large language models on educational mathematics tasks. The focus is on step-by-step reasoning, clear explanations, and pedagogically useful responses.
Dataset Structure
The dataset is provided in JSONL format. Each line contains a single instruction-response pair.
Intended Use
Fine-tuning LLMs for math tutoring
Educational… See the full description on the dataset page: https://huggingface.co/datasets/ushasree2001/educational-math-instruction-dataset.Multidisciplinary-Educational-Summaries
Knowledge Summarization Dataset
Overview
100 structured knowledge summaries across STEM, social sciences, and humanities. Features 70% Indian-centric content, 25% European perspectives, and 5% other Asian contexts for balanced representation.
Dataset Structure
{
"input": "Long-form text",
"output": {
"type": "summary",
"topic": "Subject name",
"difficulty": "beginner/intermediate/advanced",
"points": ["Key point 1", "Key point 2"]
}
}… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Multidisciplinary-Educational-Summaries.Gemini-Ultra-Educational-Conversation-Dataset-ConversationUnawareEducational-Flashcards-for-Global-Learners
1. Educational-Flashcards-for-Global-Learners/README.md
Educational Flashcards Dataset
Overview
A comprehensive collection of 100 educational flashcards covering STEM, humanities, law, arts, and cultural topics. Curated with 70% Indian content, 25% European, and 5% other Asian perspectives to promote diverse knowledge representation.
Dataset Structure
{
"input": "Text description",
"output": {
"type": "flashcards",
"topic": "Subject name"… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Educational-Flashcards-for-Global-Learners.adaption-andean-educational-prompts
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-andean_educational_prompts
This dataset contains prompt-completion pairs focused on designing contextualized educational activities, curricula, and evaluation tools for rural and indigenous communities in the Andes and Amazon regions. The content covers topics such as intercultural pedagogy, community-based projects, and the use of local materials for teaching subjects like math and art.… See the full description on the dataset page: https://huggingface.co/datasets/Oliver369X/adaption-andean-educational-prompts.eval_educational_prompteducationalProgramsbd-educational-instituteseducational-ai-agent-small-annotation-deptheducational-data-for-alpaca-finetuningEducationalext3-educational-dataseteducational-ai-agenteducational-small2
