datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.Nepali-Text-Corpus
Nepali Text Corpus
Overview
Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a
diverse range of text types, including news articles, blogs, and more, making it an invaluable
resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP)
and computational linguistics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.neBrahma-Nepali-Pretrain-Corpus
neBrahma Nepali Pretrain Corpus P2b
Dataset Summary
The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text
corpus assembled and certified for language model pretraining.
It contains 20,321,968 documents and 1.845 billion tokens of clean, verified
Devanagari Nepali text, drawn from four diverse sources and processed through an
eight-stage cleaning and quality pipeline.
This corpus serves as the training data for
neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.rejected-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.cc100-nepali
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from
CC-100, suitable for transformer pretraining.
Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Token Statistics (train split)
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.nepali-law-v2-corrected
Dataset Card for nepali-law-v2-corrected (V2)
Version 2.0.0 — a curated, audited correction of
aarajbhattarai/nepali-law-v2
(revision aa71fbe2b22310d45f86e3b429d3815817a33574).
This card describes V2. The original V1 dataset is unmodified and remains the
upstream source of truth. Every statistic here was computed from the released V2
files by the release audit pipeline (scripts/validate_release.py and the
project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.Nepali-Corpus
Nepali-Corpus
What Is This?
Everything combined—7.1 million rows of Nepali. News, Wikipedia, YouTube comments, all together. It's meant to be a solid foundation if you want to build NLP tools for Nepali.
Dataset Composition
Total rows: 7,167,456
Subset
Rows
Domain profile
Script profile
Full corpus
7,167,456
Formal + colloquial + encyclopedia + news
Devanagari, Latin, mixed
Formal subset
6,735,808
Formal/news/encyclopedia writing
Mostly Devanagari… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Corpus.nepal-section-wise-act-datasets
Nepal Section-wise Act Datasets
Dataset Description
This dataset contains section-wise legal acts and laws of Nepal, organized for easy access and analysis. It is designed to support legal research, natural language processing (NLP) tasks, and the development of legal tech applications in Nepal.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use and may result in legal action or… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-section-wise-act-datasets.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.nepali-recipes-qwen-processed
Nepali Recipes for Qwen Fine-tuning
Dataset Description
This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format.
Train Split: 900 recipes
Test Split: 327 recipes
Language: Nepali (ne)
Format: Qwen ChatML
Base Model: Qwen/Qwen2-1.5B
Dataset Structure
Data Fields
text: Full ChatML formatted prompt with answer (for training)
test_text: ChatML prompt without answer (for inference)
name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.gorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.nepali_news_textLegal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.cc100-nepali-cleaned
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text from
CC-100 suitable for transformer pretraining.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Created: 2026-04-02
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.nepal-constitution-dataset
Nepal Constitution Dataset
Dataset Description
This dataset contains the Constitution of Nepal (२०७२), organized section-wise for easy access, analysis, and use in NLP and legal tech applications. It is designed to support legal research, educational purposes, and the development of AI-driven tools for the Nepali legal system.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-constitution-dataset.nepali-proofreader
Nepali OCR Proofreading Dataset (Devanagari)
Dataset Summary
A Nepali-only (Devanagari script) text-correction dataset built for
fine-tuning a small language model (target: HimalayaGPT 0.5B) as an OCR
proofreader. Each example is a (corrupted, clean) pair: corrupted is
Nepali text with OCR/handwriting-style errors (character confusions,
missing matras, merged/split words, transposed or dropped characters),
and clean is the correct text it should map to.
The set is… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepali-proofreader.Nepali-Flow-Formal
Nepali-Flow-Formal
What's This?
This dataset has formal Nepali writing—the kind you'd find in news articles, encyclopedias, and research papers. Good for training language models on clear, well-written Nepali.
What's Inside
6,735,808 rows from three places:
IRIISNEPAL dataset (MIT license)
Nepali Wikipedia
Nepali news outlets (Kantipur, Setopati, etc.)
Mostly in Devanagari script. Formal writing—no slang or memes.
Schema
text
source
domain
script… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Flow-Formal.nepali_alpaca_multiturn
ShareGPT Conversations
This repository contains multi-turn human ↔ gpt conversations.
Splits
dineshkarki/nepali_alpaca_multiturn provides a split named train by default.
Usage
from datasets import load_dataset
ds = load_dataset("dineshkarki/nepali_alpaca_multiturn")
train = ds["train"]
Schema
Each row contains:
id: unique string
conversations: list of N messages (N ≥ 2), alternating human and gpt roles
Notes:
Conversations are lightly… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali_alpaca_multiturn.Nepali-Datasets-Reasoning-Grounding-V1Copyright 2026 Sandesh Bastola
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language… See the full description on the dataset page: https://huggingface.co/datasets/Matrix-Man-Lab/Nepali-Datasets-Reasoning-Grounding-V1.Nepali_law_articles_corpus
Nepal Legal & Government Education Corpus
Dataset Summary
This dataset is a collection of 868 explanatory legal and government-procedure
articles scraped from 11 trusted Nepalese sources — legal blogs,
law firm publications, government portals, and legal aid / human-rights
organizations. It was built as part of NyayaLM, a bilingual (Nepali-English)
legal foundation language model for Nepal, where it serves as one of the
general-education corpora used to ground the… See the full description on the dataset page: https://huggingface.co/datasets/gahann/Nepali_law_articles_corpus.nepali-summarization-datasetThis dataset was intended to to be used for finetuning the nepali text summerization task.
Feel free to contribute to this readme to add any information
nepali-corpus-v1
Nepali Mixed Corpus
This dataset contains a collection of Nepali text data aggregated for the purpose of fine-tuning Large Language Models (LLMs).
Dataset Description
This repository hosts a raw text corpus used to train the Llama-3-Nepali-Instruct models. It consists of mixed Nepali text sources designed to improve the vocabulary and semantic understanding of language models for the Nepali language.
Dataset Structure
The dataset is formatted as a standard text… See the full description on the dataset page: https://huggingface.co/datasets/paudelnirajan/nepali-corpus-v1.Nepali_News_DatasetThis dataset was collected and compiled from various Nepali news portals to support the community knowledge and research, not for profit.
Sources:
Kantipur,
Gorkhapatra, and
BBC Nepali (XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages)
Category
Count
News
36798
Sports
18767
Others(Mix)
7258
Opinion
2358
Entertainment
2144
Feature
2014
Diaspora
750
World
462
Education
188
Blog
30
Total
70769
The following task can be performed… See the full description on the dataset page: https://huggingface.co/datasets/caspro/Nepali_News_Dataset.nepal_5_law_RAG_QA
Nepal Legal QA — Bilingual RAG Fine-Tuning Dataset
A bilingual (English + Nepali) question-answering dataset built from 9 primary Nepali law texts for fine-tuning Small Language Models (SLMs) on Retrieval-Augmented Generation (RAG) tasks in the Nepal legal domain. Every answer is grounded in retrieved legal text with precise section/article citations.
Dataset Summary
Split
Total QA Pairs
Failed Chunks
Train
4,288
0
Test
550
0
Total
4,838
0
Both splits… See the full description on the dataset page: https://huggingface.co/datasets/chhatramani/nepal_5_law_RAG_QA.
