CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01himalaya-ai /cc100-nepali CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from CC-100, suitable for transformer pretraining. Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Token Statistics (train split) Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.tabulartext-generation1M<n<10M0 likes267 downloads6mo agoHugging Face02Aananda-giri /gorkhapatra-nepali-epaper Gorkhapatra Nepali E-Paper Corpus Per-article text extracted from PDF e-papers published on epaper.gorkhapatraonline.com, covering 11 newspaper slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal, loksewa, saturday, yuwamunch, gorkhapatra-125, other). Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.tabulartext-generation100K<n<1M0 likes51 downloads2mo agoHugging Face03dineshkarki /nepali-textbooks-corpus Nepali Textbooks Corpus for Grades 1-12 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 5634 Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12] Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.tabulartext-generation1K<n<10K2 likes43 downloads1y agoHugging Face04Titung /cc100-nepali-cleaned CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text from CC-100 suitable for transformer pretraining. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Created: 2026-04-02 Pipeline Unicode normalisation (NFC + ftfy) Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate) Language ID — fastText lid.176.bin, confidence ≥ 0.7 Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.tabulartext-generation1M<n<10M2 likes40 downloads6mo agoHugging Face05sabin1234 /NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET Nepali Devanagari SFT Dataset — Final Clean Release A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments. Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns. Dataset at a Glance Property Value Total rows 100,000 Total conversation messages 200,000 Human messages 100,000 GPT messages 100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.texttext-generation100K<n<1M0 likes24 downloads1mo agoHugging Face06dineshkarki /textbooks-qa-nepali Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbooks-qa-nepali") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of 2 messages: human and gpt… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbooks-qa-nepali.tabularquestion-answering1K<n<10K1 likes23 downloads1y agoHugging Face07lilgoose777 /nepal-law-commission-nepaligated ⚖️ Nepal Law Commission — Nepali Legal Corpus Dataset Summary A cleaned Nepali-language text corpus extracted from official annual reports published by the Nepal Law Commission (lawcommission.gov.np). The corpus spans fiscal years 2067/68 – 2081/82 (approximately 2010–2025), covering legal research, legislative drafting, law reform activities, and policy recommendations. Each row is a self-contained chunk of Nepali text (~300–1200 characters), filtered from mixed… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepal-law-commission-nepali.tabulartext-generation1K<n<10K0 likes9 downloads5mo agoHugging Face08lilgoose777 /mof-nepal-nepaligated 💰 Ministry of Finance Nepal — Nepali Government Finance Corpus Dataset Summary A cleaned Nepali-language text corpus extracted from official Ministry of Finance (MoF), Nepal ministry-wise progress reports published on mof.gov.np. The corpus spans fiscal years 2072/73 – 2080/81 (approximately 2015–2024), covering budget implementation, ministry-level expenditure, program progress, and financial reporting across all government ministries of Nepal. Each row is a… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/mof-nepal-nepali.tabulartext-generation1K<n<10K0 likes5 downloads5mo agoHugging Face09dineshkarki /nepali-textbooks-grade10 Nepali Textbooks Grade 10 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 1936 Grades: [10] Subjects: ['Civic_Science', 'Education', 'Health_and_Physical_Education', 'Population_Studies', 'Social_Studies', 'Sociology', 'computer_science', 'economics', 'environmental_science', 'health', 'history', 'math', 'nepali', 'optional_math', 'science', 'social'] Total chars: 5903753 Avg tokens per sample: 492… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-grade10.tabulartext-generation1K<n<10K0 likes4 downloads1y agoHugging Face10lilgoose777 /moha-nepal-nepaligated 🇳🇵 MoHA Nepal — Nepali Government Corpus Dataset Summary A cleaned Nepali-language text corpus extracted from official PDF documents published by the Ministry of Home Affairs (MoHA), Nepal (moha.gov.np). The corpus covers annual progress reports and quarterly disclosures spanning fiscal years 2076/77 – 2082/83 (approx. 2019–2026). Each row is a self-contained chunk of Nepali text (~300–1200 characters), cleaned of OCR artifacts and annotated with rich metadata including… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/moha-nepal-nepali.tabulartext-generation1K<n<10K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.