datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-news-dataset
🇳🇵 Nepali News Dataset & NLP Corpus
The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours.
Repository: thegauravgiri/nepali-news-dataset
Total Articles: 15,000+ full-text articles and growing
Update Frequency: Every 4 hours via automated GitHub Actions pipelines
Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding
License: MIT License
⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.scenesmith-example-scenes
SceneSmith Example Scenes
Project Page | Paper | Code
Example scenes generated by SceneSmith, a hierarchical agentic framework for constructing simulation-ready indoor environments from natural language prompts.
This dataset contains all scenes from the SceneSmith method (and its ablations) used in the paper evaluations. Each scene is a complete simulation-ready environment with 3D assets (including VLM-estimated physical properties), collision meshes, floor plans, and scene… See the full description on the dataset page: https://huggingface.co/datasets/nepfaff/scenesmith-example-scenes.nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.CoSER
CoSER Dataset
Overview
CoSER is a high-quality dataset for role-playing LLMs, sourced from 771 renowned novels. The dataset contains authentic multi-turn, multi-character dialogues extracted from acclaimed literary works.
Key Features
Authentic Content: Unlike synthetic datasets, CoSER extracts real dialogues from literature, maintaining high fidelity to the original works. The dialogues are inherently multi-turn and multi-character, exhibiting natural… See the full description on the dataset page: https://huggingface.co/datasets/Neph0s/CoSER.Nepali-Text-Corpus
Nepali Text Corpus
Overview
Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a
diverse range of text types, including news articles, blogs, and more, making it an invaluable
resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP)
and computational linguistics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.Wild-City
WildCity Dataset
WildCity is a real-world city-scale multimodal dataset for street-view reconstruction, simulation, and spatial intelligence. It is collected from autonomous-driving fleet logs across multiple U.S. cities and contains surround-view RGB images, LiDAR, calibration, ego and sensor poses, object annotations, semantic masks, and processed reconstruction assets.
This repository hosts the initial public release of WildCity. This version does not include the full raw… See the full description on the dataset page: https://huggingface.co/datasets/Neptune615/Wild-City.nepali-cs-asr
Nepali–English Code-Switched ASR
A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary.
v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.nepali-corpus-compilenepali_asr_datanepali_llm_datasets
Nepali LLM Datasets
This repository contains two configurations of Nepali LLM datasets:
Configurations
1. Scrapy Engine
Description: Contains data collected using a web scraping engine.
Files: [List any specific files or formats]
2. Nepberta
Description: This dataset is derived from the Nepberta project and contains cleaned data specifically related to the project. The dataset contains **cleaned text chunks of size ~50 mb ** of all… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/nepali_llm_datasets.nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.Nepali_pretraning_Corpusoag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.unjudged-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.NEPATEC3.0
National Environmental Policy Act Text Corpus (NEPATEC3.0)
Dataset Description
The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for… See the full description on the dataset page: https://huggingface.co/datasets/PNNL/NEPATEC3.0.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.muril-nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.sangraha_nepalineBrahma-Nepali-Pretrain-Corpus
neBrahma Nepali Pretrain Corpus P2b
Dataset Summary
The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text
corpus assembled and certified for language model pretraining.
It contains 20,321,968 documents and 1.845 billion tokens of clean, verified
Devanagari Nepali text, drawn from four diverse sources and processed through an
eight-stage cleaning and quality pipeline.
This corpus serves as the training data for
neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.nepali-text-corpus-64
Nepali Text Dataset
Overview
The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs,
and more, making it an invaluable resource for researchers, developers, and enthusiasts
in the fields of Natural Language Processing (NLP) and computational linguistics.
Dataset Details
Total Articles: ~6.4 million
Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.nepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency.
nepalipixel-synthetic-ocr-benchmark
NepaliPixel Benchmark Dataset Model Card
Overview
The output_benchmark directory contains a synthetic OCR benchmark dataset for Nepali (Devanagari) script generated using the Nepali Pixel pipeline. The dataset is designed for evaluation of OCR models with minimal augmentation noise and fully rendered pages.
Data
Samples: Approximately 15,000 image‑text pairs (generated with -n 15000).
Granularity: Includes all five levels – word, sentence… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepalipixel-synthetic-ocr-benchmark.nepali-audio-deepfake-datasetnepali-corpus-compilescenesmith-preprocessed-data
SceneSmith Preprocessed Data
Preprocessed 3D assets for use with SceneSmith, a VLM-agent-based system for generating physically realistic, interactive indoor scenes.
ArtVIP (Articulated Objects)
Simulation-ready articulated objects (cabinets, drawers, appliances, etc.) converted from the ArtVIP dataset. Assets have been converted from USD to Drake SDFormat using mesh-to-sim-asset, with:
Drake SDFormat (.sdf) model files with articulated joints
Visual meshes in… See the full description on the dataset page: https://huggingface.co/datasets/nepfaff/scenesmith-preprocessed-data.nepali-roman-pretrainnepse-market-data
NEPSE Market Data
Daily and tick-level data from the Nepal Stock Exchange, captured by an
open-source pipeline and validated at every layer boundary.
Daily prices reach back to 1995-07-20. Tick data starts 2026-08-19. That gap
is a property of the source, not a backlog - see Coverage below.
Tables
Path
Grain
Source
Coverage
eod_history_ext/
symbol x day
ShareSansar (second source)
1995-07-20 -> present, 653 symbols
eod_price/
symbol x day
NEPSE… See the full description on the dataset page: https://huggingface.co/datasets/AnkitSanjyal/nepse-market-data.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.Nepali-HealthChat
