datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.Nepali-HealthChatnepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.rejected-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.Nepali-Health-QAnepali-law-v2-corrected
Dataset Card for nepali-law-v2-corrected (V2)
Version 2.0.0 — a curated, audited correction of
aarajbhattarai/nepali-law-v2
(revision aa71fbe2b22310d45f86e3b429d3815817a33574).
This card describes V2. The original V1 dataset is unmodified and remains the
upstream source of truth. Every statistic here was computed from the released V2
files by the release audit pipeline (scripts/validate_release.py and the
project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.nepali-agri-fruit-instruct-wholerejected-nepali-agri-fruit-instruct-wholenepali-honorific-benchNepali-Health-Factnepali-sft-datasethermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.nepali-bias-dataset
Nepali Bias Language Dataset
Dataset Description
A synthetic dataset of Nepali sentences labeled for
bias categories including gender, religion, caste,
regional, appearance, social status, political, age,
and disability bias. Sentences were first labeled by
LLMs (ChatGPT, Grok) prompted with real Nepali news
context, then manually reviewed and corrected by human
annotators.
Dataset Summary
Split
Examples
Train
1,362
Validation
292… See the full description on the dataset page: https://huggingface.co/datasets/ios-ioe/nepali-bias-dataset.Future_Education_Nepal_MBBS_FAQ_Dataset
Future Education Nepal — MBBS FAQ Dataset (Nepali)
Overview
This dataset (future_education_nepal_mbbs_faq_nepali_sharegpt.jsonl) is a collection of 38 question-answering conversation pairs in Nepali, covering frequently asked questions about studying MBBS (medicine) in Nepal — primarily aimed at Indian students considering Nepal as a study destination. Each record is a single-turn human↔gpt exchange in ShareGPT-style format: a Nepali-language question about MBBS… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Future_Education_Nepal_MBBS_FAQ_Dataset.Major_Health_Indicator_FY_2073_74_Nepali_Statistical_QA
Major Health Indicators FY 2073/74 — Nepali Statistical Q&A Dataset
Overview
This dataset (nepali_sharegpt_suddha_nepali_final_cleaned.jsonl) is a collection of 587 instruction-following conversation pairs in Nepali, built entirely from Nepal's official Department of Health Services (DoHS) "Major Health Indicators" report for fiscal year 2073/74 (2016/17 AD). Each record is a single-turn human↔gpt exchange: a Nepali-language question asking for a specific… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Major_Health_Indicator_FY_2073_74_Nepali_Statistical_QA.educational_vedantu_FAQs_Nepali_sft_dataset
Vedantu CUET FAQs — Nepali SFT Dataset
A single-turn instruction-following (Q&A) dataset in Nepali, built from Vedantu's CUET (Common University Entrance Test) 2026 FAQ page. The dataset is formatted for supervised fine-tuning (SFT) in a conversations-style chat schema.
Summary
Records
49
Language
Nepali (ne / npi)
Script
Devanagari (Deva)
Format
JSONL, conversations (human/human turns)
Domain
Education — CUET exam FAQs
Task type… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_vedantu_FAQs_Nepali_sft_dataset.nepali_news_textNepali_Education_Budget_ShareGPT_Instruction_Dataset
Nepali Education Budget — ShareGPT Instruction Dataset
merged_edu_sharegpt_ne_serial.jsonl
A Nepali-language, single-turn, fact-based Question–Answer dataset built from official Nepal government education budget statistics (Central Bureau of Statistics). The dataset is formatted in ShareGPT conversation style and is intended for instruction-tuning / fine-tuning language models to answer factual, numeric, statistics-grounded questions in pure Nepali (Devanagari script).… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepali_Education_Budget_ShareGPT_Instruction_Dataset.drug_poisoning_nepali_sharegpt
README — drug_poisoning_nepali_sharegpt_cleaned.jsonl
This README file provides detailed information about the drug_poisoning_nepali_sharegpt_cleaned.jsonl dataset — file format, schema, source, subject matter, question pattern diversity, answer behaviour diversity, geographic/temporal coverage, and statistical analysis, all presented in tables.
1. General File Information
Detail
Value
File name
drug_poisoning_nepali_sharegpt_cleaned.jsonl
Format… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/drug_poisoning_nepali_sharegpt.Nepal_Education_Provincial_Budget_Tax_Dataset
Nepal Education, Provincial Budget & Tax Dataset (Nepali FAQ)
A Nepali-language, fact-grounded instruction-following (Q&A) dataset built from real Government of Nepal fiscal and education statistics. Every record is a single-turn human → gpt conversation pair in which a short factual question about a budget, tax, or provincial indicator is answered with the exact figure taken from the source document.
File name: merged_all_edu_prov_tax_serial.jsonl
Format: JSON Lines (one JSON… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_Education_Provincial_Budget_Tax_Dataset.Grounded_HPV_Nepali_MCQ
Grounded HPV & Cervical Cancer — Nepali (ShareGPT format)
File: grounded_hpv_cervical_cancer_nepali_sharegpt_final_v2.jsonl
Records: 24,597 · Format: JSON Lines (one JSON object per line) · Conversation schema: ShareGPT (human / gpt turns)
1. Dataset overview
This dataset is a synthetic, grounded, multiple-choice question-answering (MCQ) dataset in Nepali, built entirely around one topic: HPV (Human Papillomavirus) vaccination and cervical cancer statistics… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Nepali_MCQ.Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset
Nepal CRS Company FAQ — Contraception & Family Planning Nepali Q&A Dataset
1. Overview
This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about contraception and family planning methods — oral pills, emergency contraceptive pills, injectables (e.g. DMPA), implants, IUDs, and condoms. The content originates from Nepal CRS Company, a well-established Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset.Nepali_Financial_Services_FAQ
Nepali Financial Services FAQ Dataset (faq_nepali_merged_1to370.jsonl)
A curated, single-turn (question → answer) instruction-following dataset of 370 real-world Frequently Asked Questions, collected from 10 Nepali banking, microfinance, insurance, and fintech organizations, and standardized into pure Nepali (Devanagari script). The dataset is designed for fine-tuning, evaluating, or benchmarking LLMs on Nepali-language financial-domain question answering.
1.… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepali_Financial_Services_FAQ.ShareHub_FAQ_Nepali_Dataset
ShareHub FAQ Nepali Dataset
A Nepali-language, instruction-following (question–answer) dataset of 2,448 real frequently-asked-questions about securities listed on the Nepal Stock Exchange (NEPSE) — covering common stocks, debentures, and mutual funds — sourced from ShareHub. Every record is a single-turn human → gpt conversation in Devanagari script, and every record carries rich provenance and generation metadata.
File analyzed: sharehub_faq.jsonl
1. Quick Facts… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/ShareHub_FAQ_Nepali_Dataset.Nepali_Commercial_Bank_Financial_Indicators_QA_Dataset
Nepali Commercial Bank Financial Indicators — Grounded QA Dataset
A Nepali-language, grounded single-metric question–answering dataset built from the quarterly Key Financial Indicators of Commercial Banks published for Nepal's commercial banking sector. Each example is a single-turn human↔assistant conversation (ShareGPT / Hermes style) in which a question about one specific bank, one specific quarter, and one specific financial metric is answered with the exact value drawn from… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepali_Commercial_Bank_Financial_Indicators_QA_Dataset.Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.Nepal_Blood_Donation_FAQ
Nepal Blood Donation FAQ — Nepali Dataset
Overview
This dataset (nepal_blood_donation_faq_nepali.jsonl) is a collection of 26 instruction-following conversation pairs in Nepali, covering frequently asked questions about blood, blood groups, and blood donation. Each record is a single-turn human↔gpt exchange: a Nepali-language question followed by an informative Nepali-language answer.
The source is attributed to राष्ट्रिय रक्तसञ्चार सेवा केन्द्र ("National Blood… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_Blood_Donation_FAQ.Nepal_Education_Provincial_Budget_Tax_Miscellaneous_Statistics
Nepal Education, Provincial Budget, Tax & Miscellaneous Statistics — Instruction Dataset
A Nepali-language (with a small English subset) instruction-following (Q&A) dataset built from official Nepali government and statistical sources — covering federal/provincial education budgets, tax exemptions, trade statistics, cooperative membership, disaster losses, national accounts, and Department of Roads budget data.
File: merged_all_edu_prov_tax_misc_serial.jsonl
Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_Education_Provincial_Budget_Tax_Miscellaneous_Statistics.
