datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NepaliSentimentiNLTK_Nepali_News_DatasetNepTam-A-Nepali-Tamang-Parallel-Corpus
🧾 NepTam — A Nepali–Tamang Parallel Corpus
Dataset Summary
NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains:
20K gold-standard human-translated sentence pairs, and
80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus.
Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.NepaliCovidTweetsNepaliDevanagariSentimentAnalysis
Nepali Sentiment Dataset (Devanagari)
Dataset Summary
This dataset contains Nepali sentences in Devanagari script labeled with sentiment classes: Negative, Neutral, and Positive.
Supported Tasks
Sentiment analysis / text classification
Languages
Nepali (ne) — Devanagari script
Dataset Structure
Data Instances
Each example contains:
text: Nepali sentence (string)
label: one of {negative, neutral, positive}
Example:… See the full description on the dataset page: https://huggingface.co/datasets/Sagar32/NepaliDevanagariSentimentAnalysis.yt-nepali-movie-reviewsnepali-summarization-datasetThis dataset was intended to to be used for finetuning the nepali text summerization task.
Feel free to contribute to this readme to add any information
RoundTripOCR-nepaliPost-OCR error correction dataset (train, test and validation set) for Nepali language generated using RoundTripOCR technique.
Code: https://github.com/harshvivek14/RoundTripOCR
XLSum-nepali-summerization-datasetstsb_nepaliThe stsb_nepali dataset has been translated from stsb_multi_mt.
from datasets import load_dataset
dataset = load_dataset("syubraj/stsb_nepali")
nepali_sa
Sentiment Analysis Data for the Nepali Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Singh et al. (2020).
Data Structure:
The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models.
Citation:
@INPROCEEDINGS{9381292,
author={Singh, Oyesh Mann and Timilsina, Sandesh and Bal, Bal Krishna and Joshi, Anupam},
booktitle={2020 IEEE/ACM International Conference on Advances in Social… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/nepali_sa.nepali-meme-captions
NeMeme-CAP: Nepali Meme Captions
Dataset Summary
English-language captions generated by Google Gemini for the CHiPSAL 2026 SubtaskA Nepali Meme Datset.
The context-aware captions was generated accross the training, validation, and test splits.
Supported Tasks
Hateful Meme Classification: Predict whether the meme is non-hateful (label=0) and hateful (label=1).
Multimodal Meme Understanding: Useful as auxiliary text features or as ground-truth explanations for… See the full description on the dataset page: https://huggingface.co/datasets/Anish/nepali-meme-captions.NepaliLegalQueryParanepaliflow-romanized-nepali-to-devanagari-dataset
NepaliFlow Romanized Nepali to Devanagari Dataset
This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script.
Task
The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output.
Columns
prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari
completion: expected Nepali Devanagari output
Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.nepali_conceptnet
ConceptNet Data for the Nepali Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/nepali_conceptnet.nepalimetaphorcorpus
Nepali Metaphor Detection Dataset
Here metaphor includes all kind of figurative speech that can be interpreted diferrently while reading literal and mean differently in meaning. The AarthaAlankaars like Simile,oxymoron, paradox, juxtaposition, personification, proverbs and idioms/phrases are included
as a metaphor and thus annotated as metaphor. The classification of these subtypes is not done in dataset.
Dataset Card
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/bnabin/nepalimetaphorcorpus.nepali-tts-mos-resultsroman-nepali-gemma-finalnepali_food_500_dishes
Nepali Gastronomy: 500 Traditional and Modern Dishes
Overview
This dataset provides a comprehensive list of 500 food items from Nepal, representing the country's vast culinary landscape. It spans traditional staples, regional specialties from various ethnic groups (Newari, Tharu, Sherpa, Rai, Limbu, etc.), and modern street foods popular in urban centers.
Unlike smaller datasets, this collection includes detailed information on Primary Ingredients and Origin/Context… See the full description on the dataset page: https://huggingface.co/datasets/rajeshrai577/nepali_food_500_dishes.rakshak-nepali-toxicity-augmentedNepaliScienceVQAcc100-nepali-strictly-cleaned-devanagari-only
CC-100 Nepali — Cleaned(Devanagari Only)
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact deduplication (MD5)
Near-deduplication (char 13-gram bloom filter)
98/1/1 train/val/test split, seed 42
Usage
from datasets import load_dataset
ds = load_dataset("Basanta55/cc100-nepali-strictly-cleaned-devanagari-only")
Location_Names_in_Nepalinepali_gec_data_v3Roleplay-Nepali
RolePlay-Nepali
Roleplay-Nepali Dataset is a dataset for roleplaying in the Nepali language for the Large Language Model.
The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, see this github repo.
For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Nepali.restaurant-nepali-roman-reviewsift-nepali-v5roman-nepali-alpacanepali_gec_data_v4rakshak-nepali-toxicity-v2
