datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.nepali-agri-gov-instruct
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-agri-gov-instruct.rejected-nepali-agri-gov-instruct
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-agri-gov-instruct.Nepali-Text-Corpus
Nepali Text Corpus
Overview
Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a
diverse range of text types, including news articles, blogs, and more, making it an invaluable
resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP)
and computational linguistics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.unjudged-nepali-agri-gov-instruct
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-agri-gov-instruct.unjudged-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.neBrahma-Nepali-Pretrain-Corpus
neBrahma Nepali Pretrain Corpus P2b
Dataset Summary
The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text
corpus assembled and certified for language model pretraining.
It contains 20,321,968 documents and 1.845 billion tokens of clean, verified
Devanagari Nepali text, drawn from four diverse sources and processed through an
eight-stage cleaning and quality pipeline.
This corpus serves as the training data for
neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.rejected-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.cc100-nepali
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from
CC-100, suitable for transformer pretraining.
Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Token Statistics (train split)
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.nepali-law-v2-corrected
Dataset Card for nepali-law-v2-corrected (V2)
Version 2.0.0 — a curated, audited correction of
aarajbhattarai/nepali-law-v2
(revision aa71fbe2b22310d45f86e3b429d3815817a33574).
This card describes V2. The original V1 dataset is unmodified and remains the
upstream source of truth. Every statistic here was computed from the released V2
files by the release audit pipeline (scripts/validate_release.py and the
project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.Nepali-Corpus
Nepali-Corpus
What Is This?
Everything combined—7.1 million rows of Nepali. News, Wikipedia, YouTube comments, all together. It's meant to be a solid foundation if you want to build NLP tools for Nepali.
Dataset Composition
Total rows: 7,167,456
Subset
Rows
Domain profile
Script profile
Full corpus
7,167,456
Formal + colloquial + encyclopedia + news
Devanagari, Latin, mixed
Formal subset
6,735,808
Formal/news/encyclopedia writing
Mostly Devanagari… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Corpus.unjudged-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-fruit-rerun.nepal-section-wise-act-datasets
Nepal Section-wise Act Datasets
Dataset Description
This dataset contains section-wise legal acts and laws of Nepal, organized for easy access and analysis. It is designed to support legal research, natural language processing (NLP) tasks, and the development of legal tech applications in Nepal.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use and may result in legal action or… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-section-wise-act-datasets.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.nepali-recipes-qwen-processed
Nepali Recipes for Qwen Fine-tuning
Dataset Description
This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format.
Train Split: 900 recipes
Test Split: 327 recipes
Language: Nepali (ne)
Format: Qwen ChatML
Base Model: Qwen/Qwen2-1.5B
Dataset Structure
Data Fields
text: Full ChatML formatted prompt with answer (for training)
test_text: ChatML prompt without answer (for inference)
name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.gorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.nepali_news_textLegal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.cc100-nepali-cleaned
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text from
CC-100 suitable for transformer pretraining.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Created: 2026-04-02
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.nepal-constitution-dataset
Nepal Constitution Dataset
Dataset Description
This dataset contains the Constitution of Nepal (२०७२), organized section-wise for easy access, analysis, and use in NLP and legal tech applications. It is designed to support legal research, educational purposes, and the development of AI-driven tools for the Nepali legal system.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-constitution-dataset.nepali-proofreader
Nepali OCR Proofreading Dataset (Devanagari)
Dataset Summary
A Nepali-only (Devanagari script) text-correction dataset built for
fine-tuning a small language model (target: HimalayaGPT 0.5B) as an OCR
proofreader. Each example is a (corrupted, clean) pair: corrupted is
Nepali text with OCR/handwriting-style errors (character confusions,
missing matras, merged/split words, transposed or dropped characters),
and clean is the correct text it should map to.
The set is… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepali-proofreader.Nepali-Flow-Formal
Nepali-Flow-Formal
What's This?
This dataset has formal Nepali writing—the kind you'd find in news articles, encyclopedias, and research papers. Good for training language models on clear, well-written Nepali.
What's Inside
6,735,808 rows from three places:
IRIISNEPAL dataset (MIT license)
Nepali Wikipedia
Nepali news outlets (Kantipur, Setopati, etc.)
Mostly in Devanagari script. Formal writing—no slang or memes.
Schema
text
source
domain
script… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Flow-Formal.nepali-tamil-corpus
Nepali + Tamil Monolingual Pretraining Corpora — combined mirror
Artifacts for a two-model coursework project (Language Models and Agents, Monsoon 2026)
that trains two completely independent ~25M-parameter decoder-only Transformers from
scratch — one Tamil (Model H, higher-resource), one Nepali (Model L, lower-resource).
This is a combined mirror, not a multilingual dataset. The two corpora share no
documents, tokenizer, vocabulary, or weights, and neither model is initialised… See the full description on the dataset page: https://huggingface.co/datasets/LMA-Project-Resources-Vijay/nepali-tamil-corpus.nepali_alpaca_multiturn
ShareGPT Conversations
This repository contains multi-turn human ↔ gpt conversations.
Splits
dineshkarki/nepali_alpaca_multiturn provides a split named train by default.
Usage
from datasets import load_dataset
ds = load_dataset("dineshkarki/nepali_alpaca_multiturn")
train = ds["train"]
Schema
Each row contains:
id: unique string
conversations: list of N messages (N ≥ 2), alternating human and gpt roles
Notes:
Conversations are lightly… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali_alpaca_multiturn.
