CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads7d agoHugging Face02aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads7d agoHugging Face03IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes969 downloads1y agoHugging Face04tonibirat /neBrahma-Nepali-Pretrain-Corpus neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean, verified Devanagari Nepali text, drawn from four diverse sources and processed through an eight-stage cleaning and quality pipeline. This corpus serves as the training data for neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.texttext-generation10M<n<100M0 likes510 downloads3mo agoHugging Face05Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes489 downloads1y agoHugging Face06himalaya-ai /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.texttext-generation1K<n<10K4 likes363 downloads26d agoHugging Face07aarajbhattarai /nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.texttext-generation1K<n<10K0 likes323 downloads4d agoHugging Face08aarajbhattarai /rejected-nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.texttext-generation1K<n<10K0 likes321 downloads4d agoHugging Face09himalaya-ai /cc100-nepali CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from CC-100, suitable for transformer pretraining. Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Token Statistics (train split) Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.tabulartext-generation1M<n<10M0 likes263 downloads6mo agoHugging Face10aarajbhattarai /nepali-law-v2-corrected Dataset Card for nepali-law-v2-corrected (V2) Version 2.0.0 — a curated, audited correction of aarajbhattarai/nepali-law-v2 (revision aa71fbe2b22310d45f86e3b429d3815817a33574). This card describes V2. The original V1 dataset is unmodified and remains the upstream source of truth. Every statistic here was computed from the released V2 files by the release audit pipeline (scripts/validate_release.py and the project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.texttext-generation10K<n<100K0 likes137 downloads5d agoHugging Face11Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes134 downloads6mo agoHugging Face12Boredoom17 /Nepali-Corpus Nepali-Corpus What Is This? Everything combined—7.1 million rows of Nepali. News, Wikipedia, YouTube comments, all together. It's meant to be a solid foundation if you want to build NLP tools for Nepali. Dataset Composition Total rows: 7,167,456 Subset Rows Domain profile Script profile Full corpus 7,167,456 Formal + colloquial + encyclopedia + news Devanagari, Latin, mixed Formal subset 6,735,808 Formal/news/encyclopedia writing Mostly Devanagari… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Corpus.texttext-generation1M<n<10M4 likes120 downloads6mo agoHugging Face13ranjitraut /nepal-section-wise-act-datasets Nepal Section-wise Act Datasets Dataset Description This dataset contains section-wise legal acts and laws of Nepal, organized for easy access and analysis. It is designed to support legal research, natural language processing (NLP) tasks, and the development of legal tech applications in Nepal. Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use and may result in legal action or… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-section-wise-act-datasets.textquestion-answering100K<n<1M0 likes92 downloads2mo agoHugging Face14cloudfrm-site /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.texttext-generation1K<n<10K0 likes77 downloads25d agoHugging Face15sijanpaudel /nepali-recipes-qwen-processed Nepali Recipes for Qwen Fine-tuning Dataset Description This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format. Train Split: 900 recipes Test Split: 327 recipes Language: Nepali (ne) Format: Qwen ChatML Base Model: Qwen/Qwen2-1.5B Dataset Structure Data Fields text: Full ChatML formatted prompt with answer (for training) test_text: ChatML prompt without answer (for inference) name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.texttext-generation1K<n<10K0 likes75 downloads1y agoHugging Face16Aananda-giri /gorkhapatra-nepali-epaper Gorkhapatra Nepali E-Paper Corpus Per-article text extracted from PDF e-papers published on epaper.gorkhapatraonline.com, covering 11 newspaper slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal, loksewa, saturday, yuwamunch, gorkhapatra-125, other). Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.tabulartext-generation100K<n<1M0 likes52 downloads2mo agoHugging Face17iamsubingyawali /nepali_news_texttexttext-generation100K<n<1M0 likes47 downloads1y agoHugging Face18Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes44 downloads16d agoHugging Face19dineshkarki /nepali-textbooks-corpus Nepali Textbooks Corpus for Grades 1-12 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 5634 Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12] Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.tabulartext-generation1K<n<10K2 likes40 downloads1y agoHugging Face20Titung /cc100-nepali-cleaned CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text from CC-100 suitable for transformer pretraining. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Created: 2026-04-02 Pipeline Unicode normalisation (NFC + ftfy) Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate) Language ID — fastText lid.176.bin, confidence ≥ 0.7 Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.tabulartext-generation1M<n<10M2 likes40 downloads6mo agoHugging Face21ranjitraut /nepal-constitution-dataset Nepal Constitution Dataset Dataset Description This dataset contains the Constitution of Nepal (२०७२), organized section-wise for easy access, analysis, and use in NLP and legal tech applications. It is designed to support legal research, educational purposes, and the development of AI-driven tools for the Nepali legal system. Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-constitution-dataset.textquestion-answering1K<n<10K0 likes38 downloads2mo agoHugging Face22himalaya-ai /nepali-proofreader Nepali OCR Proofreading Dataset (Devanagari) Dataset Summary A Nepali-only (Devanagari script) text-correction dataset built for fine-tuning a small language model (target: HimalayaGPT 0.5B) as an OCR proofreader. Each example is a (corrupted, clean) pair: corrupted is Nepali text with OCR/handwriting-style errors (character confusions, missing matras, merged/split words, transposed or dropped characters), and clean is the correct text it should map to. The set is… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepali-proofreader.texttext-generation100K<n<1M0 likes36 downloads2mo agoHugging Face23Boredoom17 /Nepali-Flow-Formal Nepali-Flow-Formal What's This? This dataset has formal Nepali writing—the kind you'd find in news articles, encyclopedias, and research papers. Good for training language models on clear, well-written Nepali. What's Inside 6,735,808 rows from three places: IRIISNEPAL dataset (MIT license) Nepali Wikipedia Nepali news outlets (Kantipur, Setopati, etc.) Mostly in Devanagari script. Formal writing—no slang or memes. Schema text source domain script… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Flow-Formal.texttext-generation1M<n<10M1 likes34 downloads6mo agoHugging Face24dineshkarki /nepali_alpaca_multiturn ShareGPT Conversations This repository contains multi-turn human ↔ gpt conversations. Splits dineshkarki/nepali_alpaca_multiturn provides a split named train by default. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/nepali_alpaca_multiturn") train = ds["train"] Schema Each row contains: id: unique string conversations: list of N messages (N ≥ 2), alternating human and gpt roles Notes: Conversations are lightly… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali_alpaca_multiturn.texttext-generation100K<n<1M0 likes32 downloads1y agoHugging Face25Matrix-Man-Lab /Nepali-Datasets-Reasoning-Grounding-V1gatedCopyright 2026 Sandesh Bastola Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at http://www.apache.org/licenses/LICENSE-2.0 Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language… See the full description on the dataset page: https://huggingface.co/datasets/Matrix-Man-Lab/Nepali-Datasets-Reasoning-Grounding-V1.texttext-generationn<1K2 likes32 downloads22d agoHugging Face26gahann /Nepali_law_articles_corpus Nepal Legal & Government Education Corpus Dataset Summary This dataset is a collection of 868 explanatory legal and government-procedure articles scraped from 11 trusted Nepalese sources — legal blogs, law firm publications, government portals, and legal aid / human-rights organizations. It was built as part of NyayaLM, a bilingual (Nepali-English) legal foundation language model for Nepal, where it serves as one of the general-education corpora used to ground the… See the full description on the dataset page: https://huggingface.co/datasets/gahann/Nepali_law_articles_corpus.texttext-generationn<1K0 likes30 downloads2mo agoHugging Face27realsanjeev /nepali-summarization-datasetThis dataset was intended to to be used for finetuning the nepali text summerization task. Feel free to contribute to this readme to add any information textsummarization100K<n<1M2 likes29 downloads8mo agoHugging Face28paudelnirajan /nepali-corpus-v1 Nepali Mixed Corpus This dataset contains a collection of Nepali text data aggregated for the purpose of fine-tuning Large Language Models (LLMs). Dataset Description This repository hosts a raw text corpus used to train the Llama-3-Nepali-Instruct models. It consists of mixed Nepali text sources designed to improve the vocabulary and semantic understanding of language models for the Nepali language. Dataset Structure The dataset is formatted as a standard text… See the full description on the dataset page: https://huggingface.co/datasets/paudelnirajan/nepali-corpus-v1.texttext-generation100K<n<1M1 likes29 downloads9mo agoHugging Face29caspro /Nepali_News_DatasetThis dataset was collected and compiled from various Nepali news portals to support the community knowledge and research, not for profit. Sources: Kantipur, Gorkhapatra, and BBC Nepali (XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages) Category Count News 36798 Sports 18767 Others(Mix) 7258 Opinion 2358 Entertainment 2144 Feature 2014 Diaspora 750 World 462 Education 188 Blog 30 Total 70769 The following task can be performed… See the full description on the dataset page: https://huggingface.co/datasets/caspro/Nepali_News_Dataset.texttext-classification10K<n<100K0 likes27 downloads2y agoHugging Face30chhatramani /nepal_5_law_RAG_QA Nepal Legal QA — Bilingual RAG Fine-Tuning Dataset A bilingual (English + Nepali) question-answering dataset built from 9 primary Nepali law texts for fine-tuning Small Language Models (SLMs) on Retrieval-Augmented Generation (RAG) tasks in the Nepal legal domain. Every answer is grounded in retrieved legal text with precise section/article citations. Dataset Summary Split Total QA Pairs Failed Chunks Train 4,288 0 Test 550 0 Total 4,838 0 Both splits… See the full description on the dataset page: https://huggingface.co/datasets/chhatramani/nepal_5_law_RAG_QA.textquestion-answering1K<n<10K0 likes26 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.