sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments. Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns. Dataset at a Glance Property Value Total rows 100,000 Total conversation messages 200,000 Human messages 100,000 GPT messages 100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Language & Script
Dataset Hierarchy
The dataset contains:
- 21 domains
- 166 subdomains
- 21 categories
Domain → Subdomain Overview
अर्थशास्त्र
- Subdomains: 8
- अर्थशास्त्रका आधारभूत अवधारणा
- कर प्रणाली
- बजार
- बैंकिङ
- मुद्रा
- राष्ट्रिय आय
- विश्व अर्थतन्त्र
- व्यापार
कम्प्युटर तथा सूचना प्रविधि
- Subdomains: 8
- इन्टरनेट
- कम्प्युटर आधारभूत ज्ञान
- कृत्रिम बुद्धिमत्ता
- डाटाबेस
- प्रोग्रामिङ
- सफ्टवेयर
- साइबर सुरक्षा
- हार्डवेयर
कला तथा मनोरञ्जन
- Subdomains: 8
- कलाकार
- चित्रकला
- नाटक
- नृत्य
- मूर्तिकला
- विश्व कला
- संगीत
- सिनेमा
खगोल तथा अन्तरिक्ष
- Subdomains: 8
- अन्तरिक्ष अभियान
- अन्तरिक्ष वैज्ञानिक
- आकाशगंगा
- उपग्रह
- खगोलीय घटना
- ग्रह
- तारा
- सौर्यमण्डल
खेलकुद
- Subdomains: 8
- एथलेटिक्स
- ओलम्पिक
- क्रिकेट
- खेलकुद इतिहास
- टेनिस
- प्रसिद्ध खेलाडी
- फुटबल
- बास्केटबल
गणित
- Subdomains: 8
- अंकगणित
- ज्यामिति
- तथ्याङ्कशास्त्र
- त्रिकोणमिति
- बीजगणित
- मापन
- संख्या प्रणाली
- सम्भाव्यता
जीवविज्ञान
- Subdomains: 8
- आनुवंशिकता
- कोशिका
- जनावर
- जैव विविधता
- पारिस्थितिकी
- मानव शरीर
- वनस्पति
- सूक्ष्मजीव
नेपालको इतिहास
- Subdomains: 8
- आधुनिक नेपालको इतिहास
- नेपाल एकीकरण
- नेपालका ऐतिहासिक घटना
- प्राचीन नेपालको इतिहास
- मध्यकालीन नेपालको इतिहास
- राणाकालीन इतिहास
- लोकतान्त्रिक आन्दोलन
- शाहकालीन इतिहास
नेपालको भूगोल
- Subdomains: 8
- नेपालका जिल्ला
- नेपालका तराई क्षेत्र
- नेपालका ताल
- नेपालका नदी
- नेपालका पहाड
- नेपालका राष्ट्रिय निकुञ्ज
- नेपालका हिमाल
- नेपालको भौगोलिक विविधता
नेपालको राजनीति तथा शासन
- Subdomains: 8
- कार्यपालिका
- निर्वाचन प्रणाली
- नेपालका राजनीतिक व्यवस्था
- नेपालको संविधान
- नेपालको संसद
- न्यायपालिका
- संघीय शासन
- स्थानीय शासन
नेपालको संस्कृति
- Subdomains: 8
- नेपाली कला
- नेपाली चाडपर्व
- नेपाली जात्रा
- नेपाली नृत्य
- नेपाली परम्परा
- नेपाली भाषा
- नेपाली संगीत
- नेपाली साहित्य
भौतिकशास्त्र
- Subdomains: 8
- ऊर्जा
- गति
- ताप
- दाब
- ध्वनि
- प्रकाश
- बल
- विद्युत
रसायनशास्त्र
- Subdomains: 8
- अणु
- अम्ल र क्षार
- तत्व
- धातु र अधातु
- परमाणु
- यौगिक
- रसायनशास्त्रका आविष्कार
- रासायनिक प्रतिक्रिया
वातावरण
- Subdomains: 8
- जल प्रदूषण
- जलवायु परिवर्तन
- जैव विविधता
- प्राकृतिक स्रोत
- वन संरक्षण
- वातावरण संरक्षण
- वायु प्रदूषण
- हरितगृह प्रभाव
विज्ञान
- Subdomains: 8
- दैनिक जीवनको विज्ञान
- विज्ञानका प्रमुख व्यक्तित्व
- विज्ञानको इतिहास
- वैज्ञानिक आविष्कार
- वैज्ञानिक उपकरण
- वैज्ञानिक तथ्य
- वैज्ञानिक सिद्धान्त
- सामान्य विज्ञान
विश्व इतिहास
- Subdomains: 8
- आधुनिक विश्व इतिहास
- उपनिवेशवाद
- औद्योगिक क्रान्ति
- प्राचीन विश्व इतिहास
- मध्ययुगीन विश्व इतिहास
- विश्व युद्ध
- विश्वका ऐतिहासिक घटना
- विश्वका प्रमुख सभ्यता
विश्व भूगोल
- Subdomains: 8
- मरुभूमि
- महादेश
- महासागर
- विश्वका ताल
- विश्वका देश
- विश्वका नदी
- विश्वका पर्वत
- विश्वका राजधानी
विश्व संगठन तथा अन्तर्राष्ट्रिय सम्बन्ध
- Subdomains: 8
- अन्तर्राष्ट्रिय मुद्रा कोष
- अन्तर्राष्ट्रिय सम्बन्ध
- क्षेत्रीय संगठन
- दक्षिण एसियाली सहयोग संगठन
- विश्व बैंक
- विश्व स्वास्थ्य संगठन
- विश्वका प्रमुख संगठन
- संयुक्त राष्ट्रसंघ
विश्व सामान्य ज्ञान
- Subdomains: 8
- विश्व सामान्य ज्ञान
- विश्वका अभिलेख
- विश्वका उपनाम
- विश्वका प्रमुख तथ्य
- विश्वका प्रसिद्ध व्यक्तित्व
- विश्वका प्रसिद्ध स्थान
- विश्वका महत्वपूर्ण दिवस
- विश्वका रोचक तथ्य
साहित्य तथा भाषा
- Subdomains: 8
- कवि
- कृति
- नेपाली साहित्य
- भाषा
- लेखक
- विश्व साहित्य
- व्याकरण
- साहित्यिक विधा
स्वास्थ्य तथा पोषण
- Subdomains: 8
- खनिज पदार्थ
- पोषण
- भिटामिन
- मानव स्वास्थ्य
- रोग
- सरसफाइ
- स्वस्थ जीवनशैली
- स्वास्थ्यसम्बन्धी सामान्य ज्ञान
Domain Distribution
Subdomain Distribution
Category Distribution
Dataset Schema
Every record uses the following top-level fields:
Field descriptions
Conversation Format
Each example follows the standard two-message SFT pattern:
{
"conversations": [
{
"from": "human",
"value": "प्रश्न यहाँ हुन्छ।"
},
{
"from": "gpt",
"value": "उत्तर यहाँ हुन्छ।"
}
]
}Roles:
human— user instruction/questiongpt— assistant response
Load Demo
Pure Python
import json
with open(
"all_pure_nepali_FINAL_CLEAN_FINAL.json",
"r",
encoding="utf-8"
) as f:
data = json.load(f)
print("Rows:", len(data))
sample = data[0]
print("ID:", sample["id"])
print("Domain:", sample["domain"])
print("Subdomain:", sample["subdomain"])
print("Category:", sample["category"])
for message in sample["conversations"]:
print(f"\n{message['from'].upper()}:")
print(message["value"])Pandas
import json
import pandas as pd
with open(
"all_pure_nepali_FINAL_CLEAN_FINAL.json",
"r",
encoding="utf-8"
) as f:
data = json.load(f)
df = pd.json_normalize(data)
print("Shape:", df.shape)
display(df.head())Distribution Demo
print("Domains:", df["domain"].nunique())
print("Subdomains:", df["subdomain"].nunique())
print("Categories:", df["category"].nunique())
print("\nDomain distribution:")
print(df["domain"].value_counts())
print("\nSubdomain distribution:")
print(df["subdomain"].value_counts())
print("\nCategory distribution:")
print(df["category"].value_counts())Sample Record
{
"id": "9992ce84-199e-49d4-9434-87e83710ced0",
"conversations": [
{
"from": "human",
"value": "गोपालवंशीहरूभन्दा पहिले नेपालमा कुन जातिको शासन थियो भन्ने बारे स्पष्ट ऐतिहासिक प्रमाण छैन, तर पुराणहरूमा कसको उल्लेख पाइन्छ?\n\nक) नाग\nख) सुर\nग) अभिर\nघ) किँरात"
},
{
"from": "gpt",
"value": "उत्तर: क) नाग\nव्याख्या: पुराणहरूमा नेपालमा नागवंशीहरूको शासन रहेको उल्लेख पाइन्छ।"
}
],
"category": "नेपालको इतिहास",
"domain": "नेपालको इतिहास",
"subdomain": "प्राचीन नेपालको इतिहास",
"language": "ne",
"language_code": "npi",
"script": "Deva",
"source_model": "gemini-3.5-flash-lite",
"source": "synthetic",
"source_name": "nepali_mcq_sft",
"source_repo": "local_generation",
"source_config": "नेपालको इतिहास",
"source_split": "train",
"source_revision": "v1.0",
"source_row_id": "9992ce84-199e-49d4-9434-87e83710ced0",
"license": "Apache-2.0",
"license_tier": "permissive",
"task_type": "instruction-following",
"generation_type": "synthetic",
"condition": "model-generated",
"url": "",
"metadata_json": "{\"domain\": \"नेपालको इतिहास\", \"subdomain\": \"प्राचीन नेपालको इतिहास\", \"source_model\": \"gemini-3.5-flash-lite\", \"category\": \"नेपालको इतिहास\"}"
}Quality & Unicode Validation
The final release was rechecked at character level.
Structural validation
- 100,000 rows
- 0 duplicate IDs
- 0 missing IDs
- 0 missing conversation arrays
- 0 malformed message objects
- 0 invalid conversation roles
- 0 empty conversation values
Unicode validation
- 0 NFC normalization issues
- 0 replacement characters (`�`)
- 0 control characters
- 0 private-use Unicode characters
- 0 detected foreign-script contamination in the validated ranges
- 0 previously identified targeted corruption patterns remaining
Latin characters
A limited number of Latin characters remain because some records legitimately contain scientific, mathematical, or technical notation such as:
K2
COP26
p/q
x
y
chemical formulas
scientific symbolsThese were intentionally preserved rather than blindly removed.
Cleaning Performed
The final cleaning workflow addressed:
- Foreign Unicode-script contamination.
- Replacement-character corruption.
- Unicode normalization inconsistencies.
- Control-character contamination.
- Malformed Devanagari vowel signs and halants.
- Broken words caused by misplaced combining marks.
- Obvious accidental Latin fragments.
- Repeated corruption patterns such as malformed
th,z-z,s-देशहरू,J-E, and similar artifacts. - Duplicate-ID validation.
- Final schema and row-count validation.
Generation Metadata
Behavior Distribution
Behavior distribution is measured independently from domain, subdomain, and category.
Each record is assigned one deterministic user-intent / expected-assistant-behavior label based on its conversation text.
Response Style Distribution
The assistant-side response style is also summarized separately:
Behavior Taxonomy
Important: This is a behavioral analytics layer, not a new dataset field. It is derived from the existing conversations for analysis and reporting.
Behavior Diversity
Because this release is a 100,000-row MCQ SFT dataset, the top-level task behavior is intentionally consistent: every example asks the assistant to select and explain an answer.
However, the question-level behavioral intent is substantially more varied. A finer-grained taxonomy was applied to estimate what the user is actually asking the model to do.
Behavioral Intent Distribution
Reasoning Demand
Assistant Response Style
Diversity Summary
What this means
The dataset has strong consistency but weak top-level behavior diversity.
The dominant behavior is:
Multiple-choice question → select correct option → provide a short explanation.
Within that fixed structure, the user intent varies across:
- factual identification
- person/entity identification
- location identification
- chronology
- counting and quantity
- mathematical computation
- definitions/concepts
- causal “why” questions
- processes/methods
- classification/selection
- ordering/comparison
So this dataset has semantic/question-type diversity, but not broad assistant-behavior diversity.
For a general-purpose Nepali SFT dataset, additional behavioral families would be needed, such as:
- open-ended question answering
- step-by-step problem solving
- summarization
- rewriting
- translation
- extraction
- comparison
- planning
- clarification
- refusal/safety behavior
- conversational follow-up
- error correction
- structured output / JSON generation
Recommendation: treat this release as a high-volume MCQ instruction-following dataset, not as a fully behavior-diverse general-purpose SFT dataset.
Recommended Use
This dataset is suitable for:
- Nepali SFT experiments
- Instruction-following fine-tuning
- Nepali language-model evaluation
- Domain-specific response learning
- Devanagari-focused NLP research
- Dataset curation and quality-validation experiments
File
Primary release
all_pure_nepali_FINAL_CLEAN_FINAL.jsonLoad
import json
with open(
"all_pure_nepali_FINAL_CLEAN_FINAL.json",
"r",
encoding="utf-8"
) as f:
data = json.load(f)
print(len(data))Final Release Summary
Release status: PASS — Final Clean Dataset
This README was generated directly from the final JSON release so that the reported counts and schema match the current file.
