CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads7d agoHugging Face02aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads7d agoHugging Face03aarajbhattarai /nepali-agri-gov-instruct Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-agri-gov-instruct.text-generation0 likes1.1k downloads9m agoHugging Face04aarajbhattarai /rejected-nepali-agri-gov-instruct Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-agri-gov-instruct.text-generation0 likes1k downloads9m agoHugging Face05IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes969 downloads1y agoHugging Face06aarajbhattarai /unjudged-nepali-agri-gov-instruct Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-agri-gov-instruct.text-generation0 likes898 downloads9m agoHugging Face07aarajbhattarai /unjudged-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.text-generation0 likes531 downloads7d agoHugging Face08tonibirat /neBrahma-Nepali-Pretrain-Corpus neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean, verified Devanagari Nepali text, drawn from four diverse sources and processed through an eight-stage cleaning and quality pipeline. This corpus serves as the training data for neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.texttext-generation10M<n<100M0 likes510 downloads3mo agoHugging Face09Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes489 downloads1y agoHugging Face10himalaya-ai /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.texttext-generation1K<n<10K4 likes363 downloads26d agoHugging Face11aarajbhattarai /nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.texttext-generation1K<n<10K0 likes323 downloads4d agoHugging Face12aarajbhattarai /rejected-nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.texttext-generation1K<n<10K0 likes321 downloads4d agoHugging Face13himalaya-ai /cc100-nepali CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from CC-100, suitable for transformer pretraining. Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Token Statistics (train split) Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.tabulartext-generation1M<n<10M0 likes263 downloads6mo agoHugging Face14aarajbhattarai /nepali-law-v2-corrected Dataset Card for nepali-law-v2-corrected (V2) Version 2.0.0 — a curated, audited correction of aarajbhattarai/nepali-law-v2 (revision aa71fbe2b22310d45f86e3b429d3815817a33574). This card describes V2. The original V1 dataset is unmodified and remains the upstream source of truth. Every statistic here was computed from the released V2 files by the release audit pipeline (scripts/validate_release.py and the project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.texttext-generation10K<n<100K0 likes137 downloads5d agoHugging Face15Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes134 downloads6mo agoHugging Face16Boredoom17 /Nepali-Corpus Nepali-Corpus What Is This? Everything combined—7.1 million rows of Nepali. News, Wikipedia, YouTube comments, all together. It's meant to be a solid foundation if you want to build NLP tools for Nepali. Dataset Composition Total rows: 7,167,456 Subset Rows Domain profile Script profile Full corpus 7,167,456 Formal + colloquial + encyclopedia + news Devanagari, Latin, mixed Formal subset 6,735,808 Formal/news/encyclopedia writing Mostly Devanagari… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Corpus.texttext-generation1M<n<10M4 likes120 downloads6mo agoHugging Face17aarajbhattarai /unjudged-nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-fruit-rerun.text-generation0 likes106 downloads4d agoHugging Face18ranjitraut /nepal-section-wise-act-datasets Nepal Section-wise Act Datasets Dataset Description This dataset contains section-wise legal acts and laws of Nepal, organized for easy access and analysis. It is designed to support legal research, natural language processing (NLP) tasks, and the development of legal tech applications in Nepal. Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use and may result in legal action or… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-section-wise-act-datasets.textquestion-answering100K<n<1M0 likes92 downloads2mo agoHugging Face19cloudfrm-site /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.texttext-generation1K<n<10K0 likes77 downloads25d agoHugging Face20sijanpaudel /nepali-recipes-qwen-processed Nepali Recipes for Qwen Fine-tuning Dataset Description This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format. Train Split: 900 recipes Test Split: 327 recipes Language: Nepali (ne) Format: Qwen ChatML Base Model: Qwen/Qwen2-1.5B Dataset Structure Data Fields text: Full ChatML formatted prompt with answer (for training) test_text: ChatML prompt without answer (for inference) name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.texttext-generation1K<n<10K0 likes75 downloads1y agoHugging Face21Aananda-giri /gorkhapatra-nepali-epaper Gorkhapatra Nepali E-Paper Corpus Per-article text extracted from PDF e-papers published on epaper.gorkhapatraonline.com, covering 11 newspaper slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal, loksewa, saturday, yuwamunch, gorkhapatra-125, other). Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.tabulartext-generation100K<n<1M0 likes52 downloads2mo agoHugging Face22iamsubingyawali /nepali_news_texttexttext-generation100K<n<1M0 likes47 downloads1y agoHugging Face23Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes44 downloads16d agoHugging Face24dineshkarki /nepali-textbooks-corpus Nepali Textbooks Corpus for Grades 1-12 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 5634 Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12] Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.tabulartext-generation1K<n<10K2 likes40 downloads1y agoHugging Face25Titung /cc100-nepali-cleaned CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text from CC-100 suitable for transformer pretraining. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Created: 2026-04-02 Pipeline Unicode normalisation (NFC + ftfy) Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate) Language ID — fastText lid.176.bin, confidence ≥ 0.7 Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.tabulartext-generation1M<n<10M2 likes40 downloads6mo agoHugging Face26ranjitraut /nepal-constitution-dataset Nepal Constitution Dataset Dataset Description This dataset contains the Constitution of Nepal (२०७२), organized section-wise for easy access, analysis, and use in NLP and legal tech applications. It is designed to support legal research, educational purposes, and the development of AI-driven tools for the Nepali legal system. Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-constitution-dataset.textquestion-answering1K<n<10K0 likes38 downloads2mo agoHugging Face27himalaya-ai /nepali-proofreader Nepali OCR Proofreading Dataset (Devanagari) Dataset Summary A Nepali-only (Devanagari script) text-correction dataset built for fine-tuning a small language model (target: HimalayaGPT 0.5B) as an OCR proofreader. Each example is a (corrupted, clean) pair: corrupted is Nepali text with OCR/handwriting-style errors (character confusions, missing matras, merged/split words, transposed or dropped characters), and clean is the correct text it should map to. The set is… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepali-proofreader.texttext-generation100K<n<1M0 likes36 downloads2mo agoHugging Face28Boredoom17 /Nepali-Flow-Formal Nepali-Flow-Formal What's This? This dataset has formal Nepali writing—the kind you'd find in news articles, encyclopedias, and research papers. Good for training language models on clear, well-written Nepali. What's Inside 6,735,808 rows from three places: IRIISNEPAL dataset (MIT license) Nepali Wikipedia Nepali news outlets (Kantipur, Setopati, etc.) Mostly in Devanagari script. Formal writing—no slang or memes. Schema text source domain script… See the full description on the dataset page: https://huggingface.co/datasets/Boredoom17/Nepali-Flow-Formal.texttext-generation1M<n<10M1 likes34 downloads6mo agoHugging Face29LMA-Project-Resources-Vijay /nepali-tamil-corpus Nepali + Tamil Monolingual Pretraining Corpora — combined mirror Artifacts for a two-model coursework project (Language Models and Agents, Monsoon 2026) that trains two completely independent ~25M-parameter decoder-only Transformers from scratch — one Tamil (Model H, higher-resource), one Nepali (Model L, lower-resource). This is a combined mirror, not a multilingual dataset. The two corpora share no documents, tokenizer, vocabulary, or weights, and neither model is initialised… See the full description on the dataset page: https://huggingface.co/datasets/LMA-Project-Resources-Vijay/nepali-tamil-corpus.text-generation100B<n<1T0 likes33 downloads9d agoHugging Face30dineshkarki /nepali_alpaca_multiturn ShareGPT Conversations This repository contains multi-turn human ↔ gpt conversations. Splits dineshkarki/nepali_alpaca_multiturn provides a split named train by default. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/nepali_alpaca_multiturn") train = ds["train"] Schema Each row contains: id: unique string conversations: list of N messages (N ≥ 2), alternating human and gpt roles Notes: Conversations are lightly… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali_alpaca_multiturn.texttext-generation100K<n<1M0 likes32 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.