datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.Nepali-HealthChathermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.rejected-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.Nepali-Health-QAnepali-law-v2-corrected
Dataset Card for nepali-law-v2-corrected (V2)
Version 2.0.0 — a curated, audited correction of
aarajbhattarai/nepali-law-v2
(revision aa71fbe2b22310d45f86e3b429d3815817a33574).
This card describes V2. The original V1 dataset is unmodified and remains the
upstream source of truth. Every statistic here was computed from the released V2
files by the release audit pipeline (scripts/validate_release.py and the
project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.neptun.scraper
Data in this dataset
Docker & NPM
Scraped using crawl4ai.
The NPM and Docker data was scraped from docs.docker.com and docs.npmjs.com and processed using GPT-4 resulting in docker_documentation.jsonl and npm_documentation.jsonl.
The file training-data-v1.jsonl also includes Titanium, dockerNLcommands and docker_ps.
GitHub
Scraped using firecrawl.
The GitHub data was scraped from docs.github.com/en using firecrawl.A few pages might be missing in the… See the full description on the dataset page: https://huggingface.co/datasets/neptun-org/neptun.scraper.nepali-agri-fruit-instruct-wholenepali-honorific-benchrejected-nepali-agri-fruit-instruct-wholenepali-sft-datasetNepali-Health-FactAlpaca-Lora-GPT4-Swedish-RefinedThis is based on: https://huggingface.co/datasets/jeremyc/Alpaca-Lora-GPT4-Swedish
I've done extensive cleaning (but I'm not yet done).
This includes:
Purging erroneous and sometimes offensive generations by the translator
Fixing code instances up to row 10300. All code was botched. There may still be some html instances to fix, but at least all python should be valid.
NEPATEC2.0
National Environmental Policy Act Text Corpus (NEPATEC2.0)
Dataset Description
The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action.
Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing… See the full description on the dataset page: https://huggingface.co/datasets/PNNL/NEPATEC2.0.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.ENG_NEP_MED_PARALLEL
Dataset Card for Dataset Name
This dataset aims to be a state of art Nepali English Parallel translation in Medical Domain. Further on this dataset will be updated with more correct translations.
Dataset Details
Dataset Description
Curated by: Bibek Poudel
Language(s) (NLP): Nepali, English
Dataset Sources [optional]
Repository: https://www.kaggle.com/datasets/rxnach/nepali-health-forum-corpus-questions-and-answers
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Bibek-Poudel/ENG_NEP_MED_PARALLEL.nepali-bias-dataset
Nepali Bias Language Dataset
Dataset Description
A synthetic dataset of Nepali sentences labeled for
bias categories including gender, religion, caste,
regional, appearance, social status, political, age,
and disability bias. Sentences were first labeled by
LLMs (ChatGPT, Grok) prompted with real Nepali news
context, then manually reviewed and corrected by human
annotators.
Dataset Summary
Split
Examples
Train
1,362
Validation
292… See the full description on the dataset page: https://huggingface.co/datasets/ios-ioe/nepali-bias-dataset.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.NEPSE_Grounded_QA_Dataset
NEPSE Grounded QA Dataset
📊 Dataset Overview
NEPSE Grounded QA Dataset is a comprehensive, grounded question-answering dataset focusing on Nepal Stock Exchange (NEPSE) listed companies and financial securities. The dataset contains factually-grounded conversational pairs (human questions and AI-generated answers) with explicit source provenance and grounding information.
This dataset is specifically designed for:
Building QA systems for Nepali financial domain… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPSE_Grounded_QA_Dataset.NepDoHS_Target_Population_2072_73_MCQ_ne
🇳🇵 nepali_sharegpt_pure_final.jsonl
Nepal Health Demographics — Synthetic MCQ Instruction Dataset (ShareGPT Format)
📖 विषयसूची (Table of Contents)
परिचय (Overview)
फाइल जानकारी (File Info)
डेटा संरचना (Data Schema)
नमूना रेकर्ड (Sample Record)
स्रोत र प्रोभेनेन्स (Source & Provenance)
Behavior Distribution
Question Pattern / Type Distribution
Category (Generation Category) Distribution
Sub-domain Distribution (43 subtypes)… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NepDoHS_Target_Population_2072_73_MCQ_ne.Nepal_Banking_Statistics_Dataset_NBSD
OpenHermes Nepali Banking Dataset
📊 Dataset Overview
OpenHermes Nepali Banking Dataset is a comprehensive, grounded question-answering dataset focusing on Banking and Financial Services in Nepal. The dataset contains 585 factually-grounded Q&A pairs (human questions and AI-generated answers) covering saving account statistics across all 77 districts and 7 provinces of Nepal, sourced from official Nepal Rastra Bank (नेपाल राष्ट्र बैंक) data.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_Banking_Statistics_Dataset_NBSD.nepali_news_textFuture_Education_Nepal_MBBS_FAQ_Dataset
Future Education Nepal — MBBS FAQ Dataset (Nepali)
Overview
This dataset (future_education_nepal_mbbs_faq_nepali_sharegpt.jsonl) is a collection of 38 question-answering conversation pairs in Nepali, covering frequently asked questions about studying MBBS (medicine) in Nepal — primarily aimed at Indian students considering Nepal as a study destination. Each record is a single-turn human↔gpt exchange in ShareGPT-style format: a Nepali-language question about MBBS… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Future_Education_Nepal_MBBS_FAQ_Dataset.nepali-banana-instruct-test
Nepali Source-Grounded Instruction Dataset
Synthetic instruction-tuning dataset in Nepali, generated with NVIDIA NeMo
Data Designer from authoritative Nepali documents (agriculture manuals from
the Government of Nepal fruit development program, and legal texts). Every
answer is grounded strictly in the source documents; unanswerable questions
are answered with an explicit refusal sentence.
90 records (chat messages format + rich metadata)
15 task types: factual retrieval… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-banana-instruct-test.educational_vedantu_FAQs_Nepali_sft_dataset
Vedantu CUET FAQs — Nepali SFT Dataset
A single-turn instruction-following (Q&A) dataset in Nepali, built from Vedantu's CUET (Common University Entrance Test) 2026 FAQ page. The dataset is formatted for supervised fine-tuning (SFT) in a conversations-style chat schema.
Summary
Records
49
Language
Nepali (ne / npi)
Script
Devanagari (Deva)
Format
JSONL, conversations (human/human turns)
Domain
Education — CUET exam FAQs
Task type… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_vedantu_FAQs_Nepali_sft_dataset.nepglish-banking-dataset
NepGlish Banking NLU Dataset
3,157 synthetic NepGlish (Nepali–English code-switched) banking queries annotated with
20 intent classes and 6 slot types, designed for fine-tuning LLMs on financial NLU
tasks in the Nepali context.
Dataset Statistics
Split
Samples
Train
2,848
Validation
147
Test
162
Total
3,157
Slot types: recipient, amount (integer NPR), frequency (daily / weekly / monthly),
card_number, biller, account_type (SAVINGS / CURRENT).… See the full description on the dataset page: https://huggingface.co/datasets/nishaantshah/nepglish-banking-dataset.Major_Health_Indicator_FY_2073_74_Nepali_Statistical_QA
Major Health Indicators FY 2073/74 — Nepali Statistical Q&A Dataset
Overview
This dataset (nepali_sharegpt_suddha_nepali_final_cleaned.jsonl) is a collection of 587 instruction-following conversation pairs in Nepali, built entirely from Nepal's official Department of Health Services (DoHS) "Major Health Indicators" report for fiscal year 2073/74 (2016/17 AD). Each record is a single-turn human↔gpt exchange: a Nepali-language question asking for a specific… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Major_Health_Indicator_FY_2073_74_Nepali_Statistical_QA.drug_poisoning_nepali_sharegpt
README — drug_poisoning_nepali_sharegpt_cleaned.jsonl
This README file provides detailed information about the drug_poisoning_nepali_sharegpt_cleaned.jsonl dataset — file format, schema, source, subject matter, question pattern diversity, answer behaviour diversity, geographic/temporal coverage, and statistical analysis, all presented in tables.
1. General File Information
Detail
Value
File name
drug_poisoning_nepali_sharegpt_cleaned.jsonl
Format… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/drug_poisoning_nepali_sharegpt.
