CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Shushant /NepaliSentimenttext1K<n<10K5 likes169 downloads5y agoHugging Face02theonlysanjeev /nepal-ooc-misinformation NepOOC: Bilingual Nepali-English Out-of-Context Multimodal Misinformation Benchmark Dataset Description NepOOC is the first publicly available Nepali-dominant, bilingual benchmark for out-of-context (OOC) multimodal misinformation detection. OOC misinformation pairs an authentic, unmanipulated image with a misleading caption to construct a false narrative, without any image manipulation — making detection a problem of image-caption semantic alignment rather than… See the full description on the dataset page: https://huggingface.co/datasets/theonlysanjeev/nepal-ooc-misinformation.imagetext-classification1K<n<10K0 likes82 downloads3mo agoHugging Face03bpHigh /iNLTK_Nepali_News_Datasettext1K<n<10K2 likes51 downloads2y agoHugging Face04ilprl-docse /NepTam-A-Nepali-Tamang-Parallel-Corpus 🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.texttranslation10K<n<100K1 likes43 downloads11mo agoHugging Face05Shushant /NepaliCovidTweetstext10K<n<100K1 likes37 downloads4y agoHugging Face06Sagar32 /NepaliDevanagariSentimentAnalysis Nepali Sentiment Dataset (Devanagari) Dataset Summary This dataset contains Nepali sentences in Devanagari script labeled with sentiment classes: Negative, Neutral, and Positive. Supported Tasks Sentiment analysis / text classification Languages Nepali (ne) — Devanagari script Dataset Structure Data Instances Each example contains: text: Nepali sentence (string) label: one of {negative, neutral, positive} Example:… See the full description on the dataset page: https://huggingface.co/datasets/Sagar32/NepaliDevanagariSentimentAnalysis.texttext-classification10K<n<100K0 likes33 downloads9mo agoHugging Face07cfilt /RoundTripOCR-nepaliPost-OCR error correction dataset (train, test and validation set) for Nepali language generated using RoundTripOCR technique. Code: https://github.com/harshvivek14/RoundTripOCR text1M<n<10M3 likes31 downloads2y agoHugging Face08manirai91 /yt-nepali-movie-reviewstextn<1K0 likes30 downloads4y agoHugging Face09realsanjeev /nepali-summarization-datasetThis dataset was intended to to be used for finetuning the nepali text summerization task. Feel free to contribute to this readme to add any information textsummarization100K<n<1M2 likes29 downloads8mo agoHugging Face10realsanjeev /XLSum-nepali-summerization-datasettextsummarization1K<n<10K1 likes24 downloads3y agoHugging Face11DGurgurov /nepali_sa Sentiment Analysis Data for the Nepali Language Dataset Description: This dataset contains a sentiment analysis dataset from Singh et al. (2020). Data Structure: The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models. Citation: @INPROCEEDINGS{9381292, author={Singh, Oyesh Mann and Timilsina, Sandesh and Bal, Bal Krishna and Joshi, Anupam}, booktitle={2020 IEEE/ACM International Conference on Advances in Social… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/nepali_sa.texttext-classification1K<n<10K0 likes24 downloads2y agoHugging Face12syubraj /stsb_nepaliThe stsb_nepali dataset has been translated from stsb_multi_mt. from datasets import load_dataset dataset = load_dataset("syubraj/stsb_nepali") textsentence-similarity1K<n<10K3 likes23 downloads2y agoHugging Face13Anish /nepali-meme-captions NeMeme-CAP: Nepali Meme Captions Dataset Summary English-language captions generated by Google Gemini for the CHiPSAL 2026 SubtaskA Nepali Meme Datset. The context-aware captions was generated accross the training, validation, and test splits. Supported Tasks Hateful Meme Classification: Predict whether the meme is non-hateful (label=0) and hateful (label=1). Multimodal Meme Understanding: Useful as auxiliary text features or as ground-truth explanations for… See the full description on the dataset page: https://huggingface.co/datasets/Anish/nepali-meme-captions.text1K<n<10K0 likes23 downloads6mo agoHugging Face14cair-nepal /ai-bias-research-landscape Dataset Card: AI Bias Research Landscape Dataset Summary This dataset contains 692 curated bibliographic records of peer-reviewed and preprint publications on artificial intelligence (AI) and algorithmic bias, published between 2012 and 2026. Each record includes publication metadata (paper title, DOI, authors, author regions, affiliations, publication year, and research domain), author ORCID identifiers, and OpenAlex-derived metadata, including OpenAlex IDs… See the full description on the dataset page: https://huggingface.co/datasets/cair-nepal/ai-bias-research-landscape.tabularn<1K0 likes23 downloads3mo agoHugging Face15acostillio /NepaliLegalQueryParatext10K<n<100K0 likes20 downloads1y agoHugging Face16unicorn-s /nepal-border-sentiment-dataset Nepal Border Sentiment Dataset YouTube comments scraped from 17 Nepali news and commentary channels covering the 2026 Nepal–India border dispute, including the Prime Minister's parliamentary remarks. The dataset is labeled for 3-class sentiment (positive / neutral / negative). Files nepal_border_comments.csv — Raw scraped comments with channel and video URL nepal_comments_labelled.csv — Translated and auto-labeled (silver standard) using… See the full description on the dataset page: https://huggingface.co/datasets/unicorn-s/nepal-border-sentiment-dataset.texttext-classification1K<n<10K0 likes19 downloads3mo agoHugging Face17DGurgurov /nepali_conceptnet ConceptNet Data for the Nepali Language Dataset Description: This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub. Data Structure: The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/nepali_conceptnet.text1K<n<10K1 likes17 downloads2y agoHugging Face18manojbaniya /roman-nepali-gemma-finaltext10K<n<100K1 likes15 downloads2y agoHugging Face19bnabin /nepalimetaphorcorpus Nepali Metaphor Detection Dataset Here metaphor includes all kind of figurative speech that can be interpreted diferrently while reading literal and mean differently in meaning. The AarthaAlankaars like Simile,oxymoron, paradox, juxtaposition, personification, proverbs and idioms/phrases are included as a metaphor and thus annotated as metaphor. The classification of these subtypes is not done in dataset. Dataset Card Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/bnabin/nepalimetaphorcorpus.texttext-classification1K<n<10K0 likes15 downloads1y agoHugging Face20dipeshch71 /nepaliflow-romanized-nepali-to-devanagari-dataset NepaliFlow Romanized Nepali to Devanagari Dataset This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script. Task The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output. Columns prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari completion: expected Nepali Devanagari output Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.text10K<n<100K1 likes15 downloads3mo agoHugging Face21rajeshrai577 /nepali_food_500_dishes Nepali Gastronomy: 500 Traditional and Modern Dishes Overview This dataset provides a comprehensive list of 500 food items from Nepal, representing the country's vast culinary landscape. It spans traditional staples, regional specialties from various ethnic groups (Newari, Tharu, Sherpa, Rai, Limbu, etc.), and modern street foods popular in urban centers. Unlike smaller datasets, this collection includes detailed information on Primary Ingredients and Origin/Context… See the full description on the dataset page: https://huggingface.co/datasets/rajeshrai577/nepali_food_500_dishes.texttext-classificationn<1K0 likes14 downloads8mo agoHugging Face22biraj-bhusal /rakshak-nepali-toxicity-augmentedtabular1K<n<10K0 likes13 downloads4mo agoHugging Face23Oshara /nepali-tts-mos-resultstabular1K<n<10K0 likes13 downloads1mo agoHugging Face24krishnamishra8848 /complete_nepal_share_market_datatabular100K<n<1M0 likes11 downloads2y agoHugging Face25nabin2004 /NepaliScienceVQAtext1K<n<10K0 likes11 downloads1y agoHugging Face26krishna0-98765467 /tourist_destination_in_nepaltext10K<n<100K0 likes10 downloads3y agoHugging Face27Basanta55 /cc100-nepali-strictly-cleaned-devanagari-only CC-100 Nepali — Cleaned(Devanagari Only) Pipeline Unicode normalisation (NFC + ftfy) Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate) Language ID — fastText lid.176.bin, confidence ≥ 0.7 Exact deduplication (MD5) Near-deduplication (char 13-gram bloom filter) 98/1/1 train/val/test split, seed 42 Usage from datasets import load_dataset ds = load_dataset("Basanta55/cc100-nepali-strictly-cleaned-devanagari-only") tabular1M<n<10M0 likes10 downloads4mo agoHugging Face28AIsumit123 /nepali_gec_data_v3text1K<n<10K0 likes9 downloads7mo agoHugging Face29nabin2004 /Location_Names_in_Nepaligeospatialn<1K0 likes8 downloads1y agoHugging Face30jojo-ai-mst /Roleplay-Nepali RolePlay-Nepali Roleplay-Nepali Dataset is a dataset for roleplaying in the Nepali language for the Large Language Model. The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, see this github repo. For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Nepali.texttext-generation1K<n<10K2 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.