CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Shushant /NepaliSentimenttext1K<n<10K5 likes173 downloads5y agoHugging Face02bpHigh /iNLTK_Nepali_News_Datasettext1K<n<10K2 likes46 downloads2y agoHugging Face03ilprl-docse /NepTam-A-Nepali-Tamang-Parallel-Corpus 🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.texttranslation10K<n<100K1 likes44 downloads11mo agoHugging Face04Shushant /NepaliCovidTweetstext10K<n<100K1 likes35 downloads4y agoHugging Face05Sagar32 /NepaliDevanagariSentimentAnalysis Nepali Sentiment Dataset (Devanagari) Dataset Summary This dataset contains Nepali sentences in Devanagari script labeled with sentiment classes: Negative, Neutral, and Positive. Supported Tasks Sentiment analysis / text classification Languages Nepali (ne) — Devanagari script Dataset Structure Data Instances Each example contains: text: Nepali sentence (string) label: one of {negative, neutral, positive} Example:… See the full description on the dataset page: https://huggingface.co/datasets/Sagar32/NepaliDevanagariSentimentAnalysis.texttext-classification10K<n<100K0 likes32 downloads9mo agoHugging Face06manirai91 /yt-nepali-movie-reviewstextn<1K0 likes30 downloads4y agoHugging Face07realsanjeev /nepali-summarization-datasetThis dataset was intended to to be used for finetuning the nepali text summerization task. Feel free to contribute to this readme to add any information textsummarization100K<n<1M2 likes29 downloads8mo agoHugging Face08cfilt /RoundTripOCR-nepaliPost-OCR error correction dataset (train, test and validation set) for Nepali language generated using RoundTripOCR technique. Code: https://github.com/harshvivek14/RoundTripOCR text1M<n<10M3 likes29 downloads2y agoHugging Face09realsanjeev /XLSum-nepali-summerization-datasettextsummarization1K<n<10K1 likes25 downloads3y agoHugging Face10syubraj /stsb_nepaliThe stsb_nepali dataset has been translated from stsb_multi_mt. from datasets import load_dataset dataset = load_dataset("syubraj/stsb_nepali") textsentence-similarity1K<n<10K3 likes24 downloads2y agoHugging Face11DGurgurov /nepali_sa Sentiment Analysis Data for the Nepali Language Dataset Description: This dataset contains a sentiment analysis dataset from Singh et al. (2020). Data Structure: The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models. Citation: @INPROCEEDINGS{9381292, author={Singh, Oyesh Mann and Timilsina, Sandesh and Bal, Bal Krishna and Joshi, Anupam}, booktitle={2020 IEEE/ACM International Conference on Advances in Social… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/nepali_sa.texttext-classification1K<n<10K0 likes23 downloads2y agoHugging Face12Anish /nepali-meme-captions NeMeme-CAP: Nepali Meme Captions Dataset Summary English-language captions generated by Google Gemini for the CHiPSAL 2026 SubtaskA Nepali Meme Datset. The context-aware captions was generated accross the training, validation, and test splits. Supported Tasks Hateful Meme Classification: Predict whether the meme is non-hateful (label=0) and hateful (label=1). Multimodal Meme Understanding: Useful as auxiliary text features or as ground-truth explanations for… See the full description on the dataset page: https://huggingface.co/datasets/Anish/nepali-meme-captions.text1K<n<10K0 likes23 downloads6mo agoHugging Face13acostillio /NepaliLegalQueryParatext10K<n<100K0 likes19 downloads1y agoHugging Face14dipeshch71 /nepaliflow-romanized-nepali-to-devanagari-dataset NepaliFlow Romanized Nepali to Devanagari Dataset This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script. Task The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output. Columns prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari completion: expected Nepali Devanagari output Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.text10K<n<100K1 likes19 downloads3mo agoHugging Face15DGurgurov /nepali_conceptnet ConceptNet Data for the Nepali Language Dataset Description: This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub. Data Structure: The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/nepali_conceptnet.text1K<n<10K1 likes18 downloads2y agoHugging Face16bnabin /nepalimetaphorcorpus Nepali Metaphor Detection Dataset Here metaphor includes all kind of figurative speech that can be interpreted diferrently while reading literal and mean differently in meaning. The AarthaAlankaars like Simile,oxymoron, paradox, juxtaposition, personification, proverbs and idioms/phrases are included as a metaphor and thus annotated as metaphor. The classification of these subtypes is not done in dataset. Dataset Card Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/bnabin/nepalimetaphorcorpus.texttext-classification1K<n<10K0 likes15 downloads1y agoHugging Face17Oshara /nepali-tts-mos-resultstabular1K<n<10K0 likes15 downloads1mo agoHugging Face18manojbaniya /roman-nepali-gemma-finaltext10K<n<100K1 likes14 downloads2y agoHugging Face19rajeshrai577 /nepali_food_500_dishes Nepali Gastronomy: 500 Traditional and Modern Dishes Overview This dataset provides a comprehensive list of 500 food items from Nepal, representing the country's vast culinary landscape. It spans traditional staples, regional specialties from various ethnic groups (Newari, Tharu, Sherpa, Rai, Limbu, etc.), and modern street foods popular in urban centers. Unlike smaller datasets, this collection includes detailed information on Primary Ingredients and Origin/Context… See the full description on the dataset page: https://huggingface.co/datasets/rajeshrai577/nepali_food_500_dishes.texttext-classificationn<1K0 likes14 downloads8mo agoHugging Face20biraj-bhusal /rakshak-nepali-toxicity-augmentedtabular1K<n<10K0 likes12 downloads4mo agoHugging Face21nabin2004 /NepaliScienceVQAtext1K<n<10K0 likes11 downloads1y agoHugging Face22Basanta55 /cc100-nepali-strictly-cleaned-devanagari-only CC-100 Nepali — Cleaned(Devanagari Only) Pipeline Unicode normalisation (NFC + ftfy) Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate) Language ID — fastText lid.176.bin, confidence ≥ 0.7 Exact deduplication (MD5) Near-deduplication (char 13-gram bloom filter) 98/1/1 train/val/test split, seed 42 Usage from datasets import load_dataset ds = load_dataset("Basanta55/cc100-nepali-strictly-cleaned-devanagari-only") tabular1M<n<10M0 likes10 downloads4mo agoHugging Face23nabin2004 /Location_Names_in_Nepaligeospatialn<1K0 likes8 downloads1y agoHugging Face24AIsumit123 /nepali_gec_data_v3text1K<n<10K0 likes8 downloads7mo agoHugging Face25jojo-ai-mst /Roleplay-Nepali RolePlay-Nepali Roleplay-Nepali Dataset is a dataset for roleplaying in the Nepali language for the Large Language Model. The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, see this github repo. For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Nepali.texttext-generation1K<n<10K2 likes7 downloads2y agoHugging Face26amirpoudel /restaurant-nepali-roman-reviewstextn<1K0 likes6 downloads2y agoHugging Face27manojbaniya /ift-nepali-v5text10K<n<100K1 likes6 downloads2y agoHugging Face28manojbaniya /roman-nepali-alpacatext10K<n<100K1 likes6 downloads2y agoHugging Face29AIsumit123 /nepali_gec_data_v4text1K<n<10K0 likes6 downloads7mo agoHugging Face30biraj-bhusal /rakshak-nepali-toxicity-v2tabularn<1K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.