datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.rejected-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.nepali-law-v2-corrected
Dataset Card for nepali-law-v2-corrected (V2)
Version 2.0.0 — a curated, audited correction of
aarajbhattarai/nepali-law-v2
(revision aa71fbe2b22310d45f86e3b429d3815817a33574).
This card describes V2. The original V1 dataset is unmodified and remains the
upstream source of truth. Every statistic here was computed from the released V2
files by the release audit pipeline (scripts/validate_release.py and the
project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.nepali_news_textLegal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.nepali-banana-instruct-test
Nepali Source-Grounded Instruction Dataset
Synthetic instruction-tuning dataset in Nepali, generated with NVIDIA NeMo
Data Designer from authoritative Nepali documents (agriculture manuals from
the Government of Nepal fruit development program, and legal texts). Every
answer is grounded strictly in the source documents; unanswerable questions
are answered with an explicit refusal sentence.
90 records (chat messages format + rich metadata)
15 task types: factual retrieval… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-banana-instruct-test.nepali-environment-sft
🌿 Nepali Environment Awareness SFT Dataset
A Supervised Fine-Tuning (SFT) dataset for Nepali-language instruction-following, focused entirely on environment awareness topics. All conversations are in Devanagari script, synthetically generated using a Nepali language model, and cleaned for training-ready quality.
📋 Dataset Summary
Field
Value
Dataset Name
nepali_environment_awareness_sft
Version
v1.0 (cleaned: v2)
Language
Nepali (ne)
Script… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/nepali-environment-sft.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.nepali-apple-rootstock-instruct-test
Nepali Source-Grounded Instruction Dataset
Synthetic instruction-tuning dataset in Nepali, generated with NVIDIA NeMo
Data Designer from authoritative Nepali documents (agriculture manuals from
the Government of Nepal fruit development program, and legal texts). Every
answer is grounded strictly in the source documents; unanswerable questions
are answered with an explicit refusal sentence.
99 records (chat messages format + rich metadata)
15 task types: factual retrieval… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-apple-rootstock-instruct-test.Nepali_SFT-Dataset_behaviourdiversity
Detailed README — Nepali Economics Q/A Datasets
1. Purpose
This document records the structure, diversity, behavior, quality-cleaning history, and recommended loading procedure for the two Nepali economics Q/A JSONL datasets. Counts and distributions are derived from the uploaded files. The historical inspection counts supplied for the pre-cleaning stage are explicitly separated from the post-cleaning file statistics.
2. Dataset Files
File… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepali_SFT-Dataset_behaviourdiversity.Multilingual-Nepali-Customer-Care-Services-Datasetnepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset
🧠 Nepali Psychology Question Dataset — 2,000 Samples
📌 Overview
The Nepali Psychology Question Dataset is a specialized Nepali-language dataset containing 2,000 psychology-related question-answer records designed for Natural Language Processing (NLP), Large Language Models (LLMs), Small Language Models (SLMs), Supervised Fine-Tuning (SFT), Question Answering (QA), instruction tuning, educational AI, and psychology-domain research.
The dataset is designed with a… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/nepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset.Customer-Care-Services-Dataset-in-Nepalinepali-textbooks-grade10
Nepali Textbooks Grade 10
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 1936
Grades: [10]
Subjects: ['Civic_Science', 'Education', 'Health_and_Physical_Education', 'Population_Studies', 'Social_Studies', 'Sociology', 'computer_science', 'economics', 'environmental_science', 'health', 'history', 'math', 'nepali', 'optional_math', 'science', 'social']
Total chars: 5903753
Avg tokens per sample: 492… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-grade10.
