datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-honorific-benchnepali-bias-dataset
Nepali Bias Language Dataset
Dataset Description
A synthetic dataset of Nepali sentences labeled for
bias categories including gender, religion, caste,
regional, appearance, social status, political, age,
and disability bias. Sentences were first labeled by
LLMs (ChatGPT, Grok) prompted with real Nepali news
context, then manually reviewed and corrected by human
annotators.
Dataset Summary
Split
Examples
Train
1,362
Validation
292… See the full description on the dataset page: https://huggingface.co/datasets/ios-ioe/nepali-bias-dataset.universalml__NepaliGPT-2.0-details
Dataset Card for Evaluation run of universalml/NepaliGPT-2.0
Dataset automatically created during the evaluation run of model universalml/NepaliGPT-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/universalml__NepaliGPT-2.0-details.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.gemma4-e2b-nepali-sft-pairs
Nepali SFT pairs for Gemma 4 E2B
468 (English prompt -> Nepali answer) pairs, the exact training data behind
saliltambe/gemma-4-E2B-it-nepali-lora.
Published so the training notebook can skip a ~13 minute generation step and so anyone
reproducing it evaluates on the same held-out split.
Provenance
Prompts: English conversation openers from
OpenAssistant/oasst1 (Apache-2.0,
human-written), filtered to role == "prompter", parent_id is None, lang == "en".
Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.shivam9980__NEPALI-LLM-details
Dataset Card for Evaluation run of shivam9980/NEPALI-LLM
Dataset automatically created during the evaluation run of model shivam9980/NEPALI-LLM
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/shivam9980__NEPALI-LLM-details.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.nepali-textbooks-math-grade10
Nepali Textbook Pretraining Corpus — Sample (Grade 10 Math)
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Schema
id — unique segment id
source — book/source string
book_id — filename-derived id (if present)
subject — subject label (e.g., "math")
grade — class/grade
chapter_index — chapter number
chapter_title — chapter/unit name
segment_index — index within chapter
text — content chunk
tokens_approx — rough token count… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-math-grade10.nepali-stt-annotationsnepali-textbooks-grade10
Nepali Textbooks Grade 10
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 1936
Grades: [10]
Subjects: ['Civic_Science', 'Education', 'Health_and_Physical_Education', 'Population_Studies', 'Social_Studies', 'Sociology', 'computer_science', 'economics', 'environmental_science', 'health', 'history', 'math', 'nepali', 'optional_math', 'science', 'social']
Total chars: 5903753
Avg tokens per sample: 492… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-grade10.
