datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc100-nepali
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from
CC-100, suitable for transformer pretraining.
Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Token Statistics (train split)
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.gorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.cc100-nepali-cleaned
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text from
CC-100 suitable for transformer pretraining.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Created: 2026-04-02
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.textbooks-qa-nepali
Textbook Question-Answering Dataset (Nepali)
This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline.
Splits
train: validated conversations with non-empty question, answer, and rephrased_text.
Usage
from datasets import load_dataset
ds = load_dataset("dineshkarki/textbooks-qa-nepali")
train = ds["train"]
Schema
train: each row contains:
id: unique string
conversations: list of 2 messages: human and gpt… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbooks-qa-nepali.nepal-law-commission-nepali
⚖️ Nepal Law Commission — Nepali Legal Corpus
Dataset Summary
A cleaned Nepali-language text corpus extracted from official annual reports published by the Nepal Law Commission (lawcommission.gov.np). The corpus spans fiscal years 2067/68 – 2081/82 (approximately 2010–2025), covering legal research, legislative drafting, law reform activities, and policy recommendations.
Each row is a self-contained chunk of Nepali text (~300–1200 characters), filtered from mixed… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepal-law-commission-nepali.mof-nepal-nepali
💰 Ministry of Finance Nepal — Nepali Government Finance Corpus
Dataset Summary
A cleaned Nepali-language text corpus extracted from official Ministry of Finance (MoF), Nepal ministry-wise progress reports published on mof.gov.np. The corpus spans fiscal years 2072/73 – 2080/81 (approximately 2015–2024), covering budget implementation, ministry-level expenditure, program progress, and financial reporting across all government ministries of Nepal.
Each row is a… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/mof-nepal-nepali.nepali-textbooks-grade10
Nepali Textbooks Grade 10
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 1936
Grades: [10]
Subjects: ['Civic_Science', 'Education', 'Health_and_Physical_Education', 'Population_Studies', 'Social_Studies', 'Sociology', 'computer_science', 'economics', 'environmental_science', 'health', 'history', 'math', 'nepali', 'optional_math', 'science', 'social']
Total chars: 5903753
Avg tokens per sample: 492… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-grade10.moha-nepal-nepali
🇳🇵 MoHA Nepal — Nepali Government Corpus
Dataset Summary
A cleaned Nepali-language text corpus extracted from official PDF documents published by the Ministry of Home Affairs (MoHA), Nepal (moha.gov.np). The corpus covers annual progress reports and quarterly disclosures spanning fiscal years 2076/77 – 2082/83 (approx. 2019–2026).
Each row is a self-contained chunk of Nepali text (~300–1200 characters), cleaned of OCR artifacts and annotated with rich metadata including… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/moha-nepal-nepali.
