CoolFace
Datasetpublic

sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET

Nepali Devanagari SFT Dataset — Final Clean Release A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments. Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns. Dataset at a Glance Property Value Total rows 100,000 Total conversation messages 200,000 Human messages 100,000 GPT messages 100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes21downloads
Dataset Card

Nepali Devanagari SFT Dataset — Final Clean Release

A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.

Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.

Dataset at a Glance

PropertyValue
Total rows100,000
Total conversation messages200,000
Human messages100,000
GPT messages100,000
Unique domains21
Unique subdomains166
Unique categories21
Languages1
Scripts1
Source models1
Task types1
Generation types1
Conditions1
Licenses1

Language & Script

FieldDistribution
language{'ne': 100000}
script{'Deva': 100000}
Primary languageNepali (`ne`)
Primary scriptDevanagari (`Deva`)

Dataset Hierarchy

The dataset contains:

  • —21 domains
  • —166 subdomains
  • —21 categories

Domain → Subdomain Overview

अर्थशास्त्र

  • —Subdomains: 8
  • —अर्थशास्त्रका आधारभूत अवधारणा
  • —कर प्रणाली
  • —बजार
  • —बैंकिङ
  • —मुद्रा
  • —राष्ट्रिय आय
  • —विश्व अर्थतन्त्र
  • —व्यापार

कम्प्युटर तथा सूचना प्रविधि

  • —Subdomains: 8
  • —इन्टरनेट
  • —कम्प्युटर आधारभूत ज्ञान
  • —कृत्रिम बुद्धिमत्ता
  • —डाटाबेस
  • —प्रोग्रामिङ
  • —सफ्टवेयर
  • —साइबर सुरक्षा
  • —हार्डवेयर

कला तथा मनोरञ्जन

  • —Subdomains: 8
  • —कलाकार
  • —चित्रकला
  • —नाटक
  • —नृत्य
  • —मूर्तिकला
  • —विश्व कला
  • —संगीत
  • —सिनेमा

खगोल तथा अन्तरिक्ष

  • —Subdomains: 8
  • —अन्तरिक्ष अभियान
  • —अन्तरिक्ष वैज्ञानिक
  • —आकाशगंगा
  • —उपग्रह
  • —खगोलीय घटना
  • —ग्रह
  • —तारा
  • —सौर्यमण्डल

खेलकुद

  • —Subdomains: 8
  • —एथलेटिक्स
  • —ओलम्पिक
  • —क्रिकेट
  • —खेलकुद इतिहास
  • —टेनिस
  • —प्रसिद्ध खेलाडी
  • —फुटबल
  • —बास्केटबल

गणित

  • —Subdomains: 8
  • —अंकगणित
  • —ज्यामिति
  • —तथ्याङ्कशास्त्र
  • —त्रिकोणमिति
  • —बीजगणित
  • —मापन
  • —संख्या प्रणाली
  • —सम्भाव्यता

जीवविज्ञान

  • —Subdomains: 8
  • —आनुवंशिकता
  • —कोशिका
  • —जनावर
  • —जैव विविधता
  • —पारिस्थितिकी
  • —मानव शरीर
  • —वनस्पति
  • —सूक्ष्मजीव

नेपालको इतिहास

  • —Subdomains: 8
  • —आधुनिक नेपालको इतिहास
  • —नेपाल एकीकरण
  • —नेपालका ऐतिहासिक घटना
  • —प्राचीन नेपालको इतिहास
  • —मध्यकालीन नेपालको इतिहास
  • —राणाकालीन इतिहास
  • —लोकतान्त्रिक आन्दोलन
  • —शाहकालीन इतिहास

नेपालको भूगोल

  • —Subdomains: 8
  • —नेपालका जिल्ला
  • —नेपालका तराई क्षेत्र
  • —नेपालका ताल
  • —नेपालका नदी
  • —नेपालका पहाड
  • —नेपालका राष्ट्रिय निकुञ्ज
  • —नेपालका हिमाल
  • —नेपालको भौगोलिक विविधता

नेपालको राजनीति तथा शासन

  • —Subdomains: 8
  • —कार्यपालिका
  • —निर्वाचन प्रणाली
  • —नेपालका राजनीतिक व्यवस्था
  • —नेपालको संविधान
  • —नेपालको संसद
  • —न्यायपालिका
  • —संघीय शासन
  • —स्थानीय शासन

नेपालको संस्कृति

  • —Subdomains: 8
  • —नेपाली कला
  • —नेपाली चाडपर्व
  • —नेपाली जात्रा
  • —नेपाली नृत्य
  • —नेपाली परम्परा
  • —नेपाली भाषा
  • —नेपाली संगीत
  • —नेपाली साहित्य

भौतिकशास्त्र

  • —Subdomains: 8
  • —ऊर्जा
  • —गति
  • —ताप
  • —दाब
  • —ध्वनि
  • —प्रकाश
  • —बल
  • —विद्युत

रसायनशास्त्र

  • —Subdomains: 8
  • —अणु
  • —अम्ल र क्षार
  • —तत्व
  • —धातु र अधातु
  • —परमाणु
  • —यौगिक
  • —रसायनशास्त्रका आविष्कार
  • —रासायनिक प्रतिक्रिया

वातावरण

  • —Subdomains: 8
  • —जल प्रदूषण
  • —जलवायु परिवर्तन
  • —जैव विविधता
  • —प्राकृतिक स्रोत
  • —वन संरक्षण
  • —वातावरण संरक्षण
  • —वायु प्रदूषण
  • —हरितगृह प्रभाव

विज्ञान

  • —Subdomains: 8
  • —दैनिक जीवनको विज्ञान
  • —विज्ञानका प्रमुख व्यक्तित्व
  • —विज्ञानको इतिहास
  • —वैज्ञानिक आविष्कार
  • —वैज्ञानिक उपकरण
  • —वैज्ञानिक तथ्य
  • —वैज्ञानिक सिद्धान्त
  • —सामान्य विज्ञान

विश्व इतिहास

  • —Subdomains: 8
  • —आधुनिक विश्व इतिहास
  • —उपनिवेशवाद
  • —औद्योगिक क्रान्ति
  • —प्राचीन विश्व इतिहास
  • —मध्ययुगीन विश्व इतिहास
  • —विश्व युद्ध
  • —विश्वका ऐतिहासिक घटना
  • —विश्वका प्रमुख सभ्यता

विश्व भूगोल

  • —Subdomains: 8
  • —मरुभूमि
  • —महादेश
  • —महासागर
  • —विश्वका ताल
  • —विश्वका देश
  • —विश्वका नदी
  • —विश्वका पर्वत
  • —विश्वका राजधानी

विश्व संगठन तथा अन्तर्राष्ट्रिय सम्बन्ध

  • —Subdomains: 8
  • —अन्तर्राष्ट्रिय मुद्रा कोष
  • —अन्तर्राष्ट्रिय सम्बन्ध
  • —क्षेत्रीय संगठन
  • —दक्षिण एसियाली सहयोग संगठन
  • —विश्व बैंक
  • —विश्व स्वास्थ्य संगठन
  • —विश्वका प्रमुख संगठन
  • —संयुक्त राष्ट्रसंघ

विश्व सामान्य ज्ञान

  • —Subdomains: 8
  • —विश्व सामान्य ज्ञान
  • —विश्वका अभिलेख
  • —विश्वका उपनाम
  • —विश्वका प्रमुख तथ्य
  • —विश्वका प्रसिद्ध व्यक्तित्व
  • —विश्वका प्रसिद्ध स्थान
  • —विश्वका महत्वपूर्ण दिवस
  • —विश्वका रोचक तथ्य

साहित्य तथा भाषा

  • —Subdomains: 8
  • —कवि
  • —कृति
  • —नेपाली साहित्य
  • —भाषा
  • —लेखक
  • —विश्व साहित्य
  • —व्याकरण
  • —साहित्यिक विधा

स्वास्थ्य तथा पोषण

  • —Subdomains: 8
  • —खनिज पदार्थ
  • —पोषण
  • —भिटामिन
  • —मानव स्वास्थ्य
  • —रोग
  • —सरसफाइ
  • —स्वस्थ जीवनशैली
  • —स्वास्थ्यसम्बन्धी सामान्य ज्ञान

Domain Distribution

DomainRows
नेपालको संस्कृति7,557
वातावरण7,010
विश्व इतिहास6,778
स्वास्थ्य तथा पोषण6,678
अर्थशास्त्र5,725
विश्व भूगोल5,578
खेलकुद4,827
नेपालको राजनीति तथा शासन4,772
गणित4,771
नेपालको भूगोल4,727
कम्प्युटर तथा सूचना प्रविधि4,630
जीवविज्ञान4,355
नेपालको इतिहास4,155
विश्व सामान्य ज्ञान4,142
विश्व संगठन तथा अन्तर्राष्ट्रिय सम्बन्ध4,065
विज्ञान4,025
खगोल तथा अन्तरिक्ष3,932
कला तथा मनोरञ्जन3,212
भौतिकशास्त्र3,102
साहित्य तथा भाषा3,068
रसायनशास्त्र2,891

Subdomain Distribution

SubdomainRows
स्वस्थ जीवनशैली1,563
जैव विविधता1,350
सरसफाइ1,273
प्राकृतिक स्रोत1,260
नेपाली चाडपर्व1,224
नेपाली परम्परा1,199
नेपाली जात्रा1,162
नेपाली साहित्य1,142
जल प्रदूषण1,070
मध्ययुगीन विश्व इतिहास1,055
विश्वका प्रमुख सभ्यता1,043
भिटामिन992
उपनिवेशवाद986
अर्थशास्त्रका आधारभूत अवधारणा983
विश्वका ऐतिहासिक घटना976
विश्वका महत्वपूर्ण दिवस974
नेपाली भाषा969
नेपालको भौगोलिक विविधता952
आधुनिक विश्व इतिहास951
वातावरण संरक्षण947
औद्योगिक क्रान्ति921
विश्वका पर्वत920
इन्टरनेट899
बजार885
नेपाली नृत्य875
महासागर874
रोग861
तथ्याङ्कशास्त्र847
नेपाली कला846
नेपालका जिल्ला846
फुटबल845
सूक्ष्मजीव841
पोषण839
वायु प्रदूषण838
वन संरक्षण830
व्याकरण824
पारिस्थितिकी823
स्थानीय शासन820
राष्ट्रिय आय812
त्रिकोणमिति806
कम्प्युटर आधारभूत ज्ञान806
विश्वका राजधानी805
कर प्रणाली803
ओलम्पिक799
विश्वका देश785
क्रिकेट771
व्यापार765
विश्वका प्रसिद्ध स्थान731
शाहकालीन इतिहास728
लोकतान्त्रिक आन्दोलन725
नेपालका राष्ट्रिय निकुञ्ज724
निर्वाचन प्रणाली709
विश्व बैंक707
हरितगृह प्रभाव691
विश्वका अभिलेख688
दैनिक जीवनको विज्ञान683
सफ्टवेयर683
संख्या प्रणाली676
कार्यपालिका672
कलाकार668
कोशिका664
गति656
रसायनशास्त्रका आविष्कार646
संयुक्त राष्ट्रसंघ638
प्राचीन नेपालको इतिहास636
मुद्रा634
बास्केटबल607
नेपालका राजनीतिक व्यवस्था604
अन्तरिक्ष अभियान599
ज्यामिति594
विश्व स्वास्थ्य संगठन592
वैज्ञानिक उपकरण592
बैंकिङ589
विज्ञानका प्रमुख व्यक्तित्व585
विश्वका ताल581
आकाशगंगा578
अन्तर्राष्ट्रिय मुद्रा कोष569
विश्वका नदी569
खेलकुद इतिहास566
नेपाली संगीत560
वैज्ञानिक आविष्कार558
आनुवंशिकता553
ग्रह547
प्रोग्रामिङ543
मापन538
न्यायपालिका537
नेपालको संविधान534
नेपालका पहाड531
अंकगणित529
खगोलीय घटना528
मरुभूमि526
एथलेटिक्स526
मानव शरीर520
खनिज पदार्थ518
महादेश518
कृत्रिम बुद्धिमत्ता515
मध्यकालीन नेपालको इतिहास515
दक्षिण एसियाली सहयोग संगठन511
राणाकालीन इतिहास506
विद्युत505
नेपाल एकीकरण497
जलवायु परिवर्तन490
नृत्य486
नेपालको संसद482
उपग्रह478
नेपालका ताल474
प्राचीन विश्व इतिहास473
मूर्तिकला468
अन्तरिक्ष वैज्ञानिक465
विश्वका प्रमुख संगठन461
विश्वका उपनाम454
बल447
ध्वनि440
सम्भाव्यता439
नेपालका हिमाल431
विश्वका प्रसिद्ध व्यक्तित्व431
विज्ञानको इतिहास429
हार्डवेयर427
टेनिस426
अणु421
तारा420
विश्व साहित्य417
संघीय शासन414
विश्व सामान्य ज्ञान414
नेपालका तराई क्षेत्र414
वैज्ञानिक तथ्य410
सिनेमा400
सामान्य विज्ञान399
डाटाबेस382
साइबर सुरक्षा375
विश्व युद्ध373
वैज्ञानिक सिद्धान्त369
रासायनिक प्रतिक्रिया358
नेपालका नदी355
धातु र अधातु348
बीजगणित342
संगीत335
स्वास्थ्यसम्बन्धी सामान्य ज्ञान329
कवि319
भाषा319
अन्तर्राष्ट्रिय सम्बन्ध318
सौर्यमण्डल317
तत्व314
प्रकाश311
मानव स्वास्थ्य303
वनस्पति297
ताप296
चित्रकला294
प्रसिद्ध खेलाडी287
नाटक281
विश्व कला280
परमाणु278
नेपालका ऐतिहासिक घटना278
विश्वका रोचक तथ्य277
आधुनिक नेपालको इतिहास270
कृति270
क्षेत्रीय संगठन269
यौगिक269
साहित्यिक विधा263
अम्ल र क्षार257
विश्व अर्थतन्त्र254
लेखक236
दाब229
ऊर्जा218
जनावर191
विश्वका प्रमुख तथ्य173

Category Distribution

CategoryRows
नेपालको संस्कृति7,557
वातावरण7,010
विश्व इतिहास6,778
स्वास्थ्य तथा पोषण6,678
अर्थशास्त्र5,725
विश्व भूगोल5,578
खेलकुद4,827
नेपालको राजनीति तथा शासन4,772
गणित4,771
नेपालको भूगोल4,727
कम्प्युटर तथा सूचना प्रविधि4,630
जीवविज्ञान4,355
नेपालको इतिहास4,155
विश्व सामान्य ज्ञान4,142
विश्व संगठन तथा अन्तर्राष्ट्रिय सम्बन्ध4,065
विज्ञान4,025
खगोल तथा अन्तरिक्ष3,932
कला तथा मनोरञ्जन3,212
भौतिकशास्त्र3,102
साहित्य तथा भाषा3,068
रसायनशास्त्र2,891

Dataset Schema

Every record uses the following top-level fields:

FieldPresent
idYes
conversationsYes
categoryYes
domainYes
subdomainYes
languageYes
language_codeYes
scriptYes
source_modelYes
sourceYes
source_nameYes
source_repoYes
source_configYes
source_splitYes
source_revisionYes
source_row_idYes
licenseYes
license_tierYes
task_typeYes
generation_typeYes
conditionYes
url100,000 missing
metadata_jsonYes

Field descriptions

FieldDescription
idUnique record identifier
conversationsUser/assistant dialogue used for SFT
categoryContent category
domainPrimary subject domain
subdomainMore specific subject area
languageLanguage identifier
language_codeSpecific language code
scriptWriting system
source_modelModel used during generation
sourceSource type
source_nameSource/dataset name
source_repoSource repository
source_configGeneration/configuration label
source_splitSplit name
source_revisionDataset revision
source_row_idSource record identifier
licenseDataset license
license_tierLicense classification
task_typeTask category
generation_typeGeneration method
conditionGeneration condition
urlSource URL, when available
metadata_jsonAdditional metadata stored as a JSON string

Conversation Format

Each example follows the standard two-message SFT pattern:

json
{
  "conversations": [
    {
      "from": "human",
      "value": "प्रश्न यहाँ हुन्छ।"
    },
    {
      "from": "gpt",
      "value": "उत्तर यहाँ हुन्छ।"
    }
  ]
}

Roles:

  • —human — user instruction/question
  • —gpt — assistant response

Load Demo

Pure Python

python
import json

with open(
    "all_pure_nepali_FINAL_CLEAN_FINAL.json",
    "r",
    encoding="utf-8"
) as f:
    data = json.load(f)

print("Rows:", len(data))

sample = data[0]

print("ID:", sample["id"])
print("Domain:", sample["domain"])
print("Subdomain:", sample["subdomain"])
print("Category:", sample["category"])

for message in sample["conversations"]:
    print(f"\n{message['from'].upper()}:")
    print(message["value"])

Pandas

python
import json
import pandas as pd

with open(
    "all_pure_nepali_FINAL_CLEAN_FINAL.json",
    "r",
    encoding="utf-8"
) as f:
    data = json.load(f)

df = pd.json_normalize(data)

print("Shape:", df.shape)
display(df.head())

Distribution Demo

python
print("Domains:", df["domain"].nunique())
print("Subdomains:", df["subdomain"].nunique())
print("Categories:", df["category"].nunique())

print("\nDomain distribution:")
print(df["domain"].value_counts())

print("\nSubdomain distribution:")
print(df["subdomain"].value_counts())

print("\nCategory distribution:")
print(df["category"].value_counts())

Sample Record

json
{
  "id": "9992ce84-199e-49d4-9434-87e83710ced0",
  "conversations": [
    {
      "from": "human",
      "value": "गोपालवंशीहरूभन्दा पहिले नेपालमा कुन जातिको शासन थियो भन्ने बारे स्पष्ट ऐतिहासिक प्रमाण छैन, तर पुराणहरूमा कसको उल्लेख पाइन्छ?\n\nक) नाग\nख) सुर\nग) अभिर\nघ) किँरात"
    },
    {
      "from": "gpt",
      "value": "उत्तर: क) नाग\nव्याख्या: पुराणहरूमा नेपालमा नागवंशीहरूको शासन रहेको उल्लेख पाइन्छ।"
    }
  ],
  "category": "नेपालको इतिहास",
  "domain": "नेपालको इतिहास",
  "subdomain": "प्राचीन नेपालको इतिहास",
  "language": "ne",
  "language_code": "npi",
  "script": "Deva",
  "source_model": "gemini-3.5-flash-lite",
  "source": "synthetic",
  "source_name": "nepali_mcq_sft",
  "source_repo": "local_generation",
  "source_config": "नेपालको इतिहास",
  "source_split": "train",
  "source_revision": "v1.0",
  "source_row_id": "9992ce84-199e-49d4-9434-87e83710ced0",
  "license": "Apache-2.0",
  "license_tier": "permissive",
  "task_type": "instruction-following",
  "generation_type": "synthetic",
  "condition": "model-generated",
  "url": "",
  "metadata_json": "{\"domain\": \"नेपालको इतिहास\", \"subdomain\": \"प्राचीन नेपालको इतिहास\", \"source_model\": \"gemini-3.5-flash-lite\", \"category\": \"नेपालको इतिहास\"}"
}

Quality & Unicode Validation

The final release was rechecked at character level.

Structural validation

  • —100,000 rows
  • —0 duplicate IDs
  • —0 missing IDs
  • —0 missing conversation arrays
  • —0 malformed message objects
  • —0 invalid conversation roles
  • —0 empty conversation values

Unicode validation

  • —0 NFC normalization issues
  • —0 replacement characters (`�`)
  • —0 control characters
  • —0 private-use Unicode characters
  • —0 detected foreign-script contamination in the validated ranges
  • —0 previously identified targeted corruption patterns remaining

Latin characters

A limited number of Latin characters remain because some records legitimately contain scientific, mathematical, or technical notation such as:

text
K2
COP26
p/q
x
y
chemical formulas
scientific symbols

These were intentionally preserved rather than blindly removed.


Cleaning Performed

The final cleaning workflow addressed:

  1. 1.Foreign Unicode-script contamination.
  2. 2.Replacement-character corruption.
  3. 3.Unicode normalization inconsistencies.
  4. 4.Control-character contamination.
  5. 5.Malformed Devanagari vowel signs and halants.
  6. 6.Broken words caused by misplaced combining marks.
  7. 7.Obvious accidental Latin fragments.
  8. 8.Repeated corruption patterns such as malformed th, z-z, s-देशहरू, J-E, and similar artifacts.
  9. 9.Duplicate-ID validation.
  10. 10.Final schema and row-count validation.

Generation Metadata

FieldDistribution
Source model{'gemini-3.5-flash-lite': 100000}
Generation type{'synthetic': 100000}
Condition{'model-generated': 100000}
Task type{'instruction-following': 100000}
License{'Apache-2.0': 100000}

Behavior Distribution

Behavior distribution is measured independently from domain, subdomain, and category.

Each record is assigned one deterministic user-intent / expected-assistant-behavior label based on its conversation text.

BehaviorRowsShare
multiple_choice_answering100,000100.00%

Response Style Distribution

The assistant-side response style is also summarized separately:

Response styleRowsShare
answer_with_explanation100,000100.00%

Behavior Taxonomy

BehaviorMeaning
multiple_choice_answeringSelects the correct option from a multiple-choice question
factual_question_answeringGives a direct fact, person, place, event, date, or count
definition_identificationDefines, identifies, or explains what a term/concept means
explanation_reasoningProvides causal or explanatory reasoning
calculation_quantitativeAnswers numeric, mathematical, ratio, percentage, or probability questions
how_to_procedureExplains a procedure, process, or method
verification_judgmentDetermines whether a statement/condition is correct
general_instruction_followingGeneral instruction-following not captured by the other buckets
unknownMissing or insufficient conversation text
Important: This is a behavioral analytics layer, not a new dataset field. It is derived from the existing conversations for analysis and reporting.

Behavior Diversity

Because this release is a 100,000-row MCQ SFT dataset, the top-level task behavior is intentionally consistent: every example asks the assistant to select and explain an answer.

However, the question-level behavioral intent is substantially more varied. A finer-grained taxonomy was applied to estimate what the user is actually asking the model to do.

Behavioral Intent Distribution

Behavioral intentRowsShare
general_mcq_factual73,69573.69%
mathematical_quantitative12,27012.27%
quantity_counting3,2663.27%
person_entity_identification2,7002.70%
causal_why_reasoning2,5472.55%
location_place_identification1,4341.43%
ordering_comparison1,3501.35%
chronology_time9260.93%
verification_judgment5100.51%
process_method4920.49%
classification_selection4670.47%
definition_concept3430.34%

Reasoning Demand

Reasoning profileRowsShare
mixed81,63781.64%
reasoning_or_computation15,97515.97%
retrieval_identification2,3882.39%

Assistant Response Style

Response styleRowsShare
answer_plus_explanation100,000100.00%

Diversity Summary

MetricValue
Observed behavioral intent classes12
Shannon entropy1.508 bits
Normalized behavioral entropy0.421
Top-level task diversityLow
Question-level intent diversityModerate
Response-style diversityLow

What this means

The dataset has strong consistency but weak top-level behavior diversity.

The dominant behavior is:

Multiple-choice question → select correct option → provide a short explanation.

Within that fixed structure, the user intent varies across:

  • —factual identification
  • —person/entity identification
  • —location identification
  • —chronology
  • —counting and quantity
  • —mathematical computation
  • —definitions/concepts
  • —causal “why” questions
  • —processes/methods
  • —classification/selection
  • —ordering/comparison

So this dataset has semantic/question-type diversity, but not broad assistant-behavior diversity.

For a general-purpose Nepali SFT dataset, additional behavioral families would be needed, such as:

  • —open-ended question answering
  • —step-by-step problem solving
  • —summarization
  • —rewriting
  • —translation
  • —extraction
  • —comparison
  • —planning
  • —clarification
  • —refusal/safety behavior
  • —conversational follow-up
  • —error correction
  • —structured output / JSON generation
Recommendation: treat this release as a high-volume MCQ instruction-following dataset, not as a fully behavior-diverse general-purpose SFT dataset.

Recommended Use

This dataset is suitable for:

  • —Nepali SFT experiments
  • —Instruction-following fine-tuning
  • —Nepali language-model evaluation
  • —Domain-specific response learning
  • —Devanagari-focused NLP research
  • —Dataset curation and quality-validation experiments

File

Primary release

text
all_pure_nepali_FINAL_CLEAN_FINAL.json

Load

python
import json

with open(
    "all_pure_nepali_FINAL_CLEAN_FINAL.json",
    "r",
    encoding="utf-8"
) as f:
    data = json.load(f)

print(len(data))

Final Release Summary

MetricValue
Rows100,000
Domains21
Subdomains166
Categories21
Duplicate IDs0
Foreign Unicode contamination0 detected
NFC issues0
Replacement characters0
Control characters0
Targeted corruption remaining0

Release status: PASS — Final Clean Dataset

This README was generated directly from the final JSON release so that the reported counts and schema match the current file.