datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.cc100-nepali
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from
CC-100, suitable for transformer pretraining.
Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Token Statistics (train split)
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.nepali-tokenizer-corpusipfs_nepal_laws_ir
Nepal legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_nepal_laws (revision 22395665af98b03f562994b5b7aca768b7e34b54) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Nepal prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_nepal_laws_ir.nepal_flood_2026
Nepal Flood 2026, Upper Trishuli and Bhote Koshi Corridor
Building inventory for the 26 August 2026 flash flood on the Nepal-China border.
AOI: 1 km buffer around the Bhote Koshi and Trishuli river centrelines, 132.93 km² across Nuwakot and Rasuwa districts. Tasking Manager project 62904.
Contents
upperstream/buildings.geojson (+ .parquet): 13,663 footprints
upperstream/building_density_h3_r8.geojson (+ .parquet): buildings per H3 res 8 cell (~0.7 km²)… See the full description on the dataset page: https://huggingface.co/datasets/hotosm/nepal_flood_2026.nepali-honorific-benchnepali-bias-dataset
Nepali Bias Language Dataset
Dataset Description
A synthetic dataset of Nepali sentences labeled for
bias categories including gender, religion, caste,
regional, appearance, social status, political, age,
and disability bias. Sentences were first labeled by
LLMs (ChatGPT, Grok) prompted with real Nepali news
context, then manually reviewed and corrected by human
annotators.
Dataset Summary
Split
Examples
Train
1,362
Validation
292… See the full description on the dataset page: https://huggingface.co/datasets/ios-ioe/nepali-bias-dataset.nepali-tokenizer-corpusaibharat_nepali_dataset_finalgorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.universalml__NepaliGPT-2.0-details
Dataset Card for Evaluation run of universalml/NepaliGPT-2.0
Dataset automatically created during the evaluation run of model universalml/NepaliGPT-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/universalml__NepaliGPT-2.0-details.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.nepali-tokenizer-corpuscc100-nepali-cleaned
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text from
CC-100 suitable for transformer pretraining.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Created: 2026-04-02
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.rakshak-nepali-toxicity-final
🛡️ RakshakAI — Augmented Nepali Toxicity Dataset
The full augmented training dataset used to train the RakshakAI toxicity detection models. Contains 4,716 samples expanded from the curated 1,574 sample dataset through back-translation augmentation via English and Hindi as intermediate languages. For the clean curated dataset only, see rakshak-all-data-combined.
📄 Paper: RakshakAI: Multi-Label Toxicity Detection for Low-Resource Nepali Social Media Content
Why this… See the full description on the dataset page: https://huggingface.co/datasets/biraj-bhusal/rakshak-nepali-toxicity-final.gemma4-e2b-nepali-sft-pairs
Nepali SFT pairs for Gemma 4 E2B
468 (English prompt -> Nepali answer) pairs, the exact training data behind
saliltambe/gemma-4-E2B-it-nepali-lora.
Published so the training notebook can skip a ~13 minute generation step and so anyone
reproducing it evaluates on the same held-out split.
Provenance
Prompts: English conversation openers from
OpenAssistant/oasst1 (Apache-2.0,
human-written), filtered to role == "prompter", parent_id is None, lang == "en".
Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.nepali-redditDataset containing posts and comments from 3 subreddits NepalSocial, technepal and nepalstock.
nepali-oov-distilled
Nepali OOV-distilled subset (854 h)
An OOV-dense distillation of Premal-12/c9nepali-audio-dataset2 (used with the
author's permission), shipped in four variants: the original single-voice audio,
a CPU-augmented copy, and 244 h re-rendered onto 1,842 real human speakers with
Seed-VC. For Nepali ASR and TTS work.
Filter with the variant field -- see Composition below. If you came here
for speaker diversity, you want variant == "vc".
What this is
The source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-oov-distilled.nepali-asr-benchmark
Nepali ASR Benchmark
Per-utterance reference, hypothesis, WER, and CER for the six released Nepali ASR
checkpoints evaluated on three independent test sets. Released alongside the paper
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech
Recognition.
Contents
Field
Type
Description
utterance_id
string
stable identifier {test_set}-{index}
reference
string
NFC-normalised gold transcription (Devanagari)
hypothesis
string… See the full description on the dataset page: https://huggingface.co/datasets/sumanpaudel1997/nepali-asr-benchmark.50k_nepali_dataset_chatbotnepal-supremecourt-judgments-cleanedsmall-newscorpus-for-nepali-ged-fullstop-removedai-bias-research-landscape
Dataset Card: AI Bias Research Landscape
Dataset Summary
This dataset contains 692 curated bibliographic records of peer-reviewed and
preprint publications on artificial intelligence (AI) and algorithmic bias,
published between 2012 and 2026. Each record includes publication metadata
(paper title, DOI, authors, author regions, affiliations, publication year,
and research domain), author ORCID identifiers, and OpenAlex-derived
metadata, including OpenAlex IDs… See the full description on the dataset page: https://huggingface.co/datasets/cair-nepal/ai-bias-research-landscape.asia-aid-flows-financial-tracking-private-sector-nepal2
Financial tracking of private sector contributions Nepal 2015
Publisher: OCHA HQ · Source: HDX · License: cc-by-igo · Updated: 2023-05-02
Abstract
Information on the private sector cash and in-kind contributions to humanitarian relief efforts in Nepal earthquake.
Each row in this dataset represents tabular records. Temporal coverage is indicated by the unnamed_11, unnamed_12 column(s). Geographic scope: NPL, NEPAL-EARTHQUAKE.
Curated into ML-ready Parquet format by… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-aid-flows-financial-tracking-private-sector-nepal2.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.textbooks-qa-nepali
Textbook Question-Answering Dataset (Nepali)
This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline.
Splits
train: validated conversations with non-empty question, answer, and rephrased_text.
Usage
from datasets import load_dataset
ds = load_dataset("dineshkarki/textbooks-qa-nepali")
train = ds["train"]
Schema
train: each row contains:
id: unique string
conversations: list of 2 messages: human and gpt… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbooks-qa-nepali.shivam9980__NEPALI-LLM-details
Dataset Card for Evaluation run of shivam9980/NEPALI-LLM
Dataset automatically created during the evaluation run of model shivam9980/NEPALI-LLM
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/shivam9980__NEPALI-LLM-details.aibharat_nepali_datasetsmall-newscorpus-for-nepali-gednepali-celeb-faces
Usage
from datasets import load_dataset
ds = load_dataset("Titung/nepali-celeb-faces")
