datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-news-dataset
🇳🇵 Nepali News Dataset & NLP Corpus
The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours.
Repository: thegauravgiri/nepali-news-dataset
Total Articles: 15,000+ full-text articles and growing
Update Frequency: Every 4 hours via automated GitHub Actions pipelines
Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding
License: MIT License
⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.Nepali-Text-Corpus
Nepali Text Corpus
Overview
Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a
diverse range of text types, including news articles, blogs, and more, making it an invaluable
resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP)
and computational linguistics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.nepali-cs-asr
Nepali–English Code-Switched ASR
A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary.
v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.nepali-corpus-compilenepali-agri-gov-instruct
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-agri-gov-instruct.rejected-nepali-agri-gov-instruct
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-agri-gov-instruct.unjudged-nepali-agri-gov-instruct
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-agri-gov-instruct.Nepali_pretraning_Corpusnepali_llm_datasets
Nepali LLM Datasets
This repository contains two configurations of Nepali LLM datasets:
Configurations
1. Scrapy Engine
Description: Contains data collected using a web scraping engine.
Files: [List any specific files or formats]
2. Nepberta
Description: This dataset is derived from the Nepberta project and contains cleaned data specifically related to the project. The dataset contains **cleaned text chunks of size ~50 mb ** of all… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/nepali_llm_datasets.unjudged-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.nepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency.
neBrahma-Nepali-Pretrain-Corpus
neBrahma Nepali Pretrain Corpus P2b
Dataset Summary
The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text
corpus assembled and certified for language model pretraining.
It contains 20,321,968 documents and 1.845 billion tokens of clean, verified
Devanagari Nepali text, drawn from four diverse sources and processed through an
eight-stage cleaning and quality pipeline.
This corpus serves as the training data for
neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.sangraha_nepalimuril-nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.nepali-text-corpus-64
Nepali Text Dataset
Overview
The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs,
and more, making it an invaluable resource for researchers, developers, and enthusiasts
in the fields of Natural Language Processing (NLP) and computational linguistics.
Dataset Details
Total Articles: ~6.4 million
Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.nepali_asr_datanepalipixel-synthetic-ocr-benchmark
NepaliPixel Benchmark Dataset Model Card
Overview
The output_benchmark directory contains a synthetic OCR benchmark dataset for Nepali (Devanagari) script generated using the Nepali Pixel pipeline. The dataset is designed for evaluation of OCR models with minimal augmentation noise and fully rendered pages.
Data
Samples: Approximately 15,000 image‑text pairs (generated with -n 15000).
Granularity: Includes all five levels – word, sentence… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepalipixel-synthetic-ocr-benchmark.nepali-audio-deepfake-datasetnepali-corpus-compilenepali-roman-pretrainnepali-kokoro-ft-data
Nepali Kokoro Fine-Tuning Dataset
This is a sharded, processed dataset containing Nepali voice data for Kokoro TTS fine-tuning.
hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.Nepali-HealthChatnepali_dataset_llmnepali-audio-reserve-r6
Nepali two-speaker conversation chunks
~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.
