CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01damo-da /oag-nepal-audit-reports OAG Nepal Audit Reports — Nepali transcripts and ruled tables Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements. The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.documenttext-retrieval10M<n<100M0 likes578 downloads23d agoHugging Face02himalaya-ai /cc100-nepali CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from CC-100, suitable for transformer pretraining. Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Token Statistics (train split) Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.tabulartext-generation1M<n<10M0 likes267 downloads6mo agoHugging Face03himalaya-ai /nepali-tokenizer-corpustabular1M<n<10M0 likes177 downloads6mo agoHugging Face04justicedao /ipfs_nepal_laws_ir Nepal legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_nepal_laws (revision 22395665af98b03f562994b5b7aca768b7e34b54) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Nepal prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_nepal_laws_ir.tabulartext-retrieval1K<n<10K0 likes135 downloads2d agoHugging Face05hotosm /nepal_flood_2026 Nepal Flood 2026, Upper Trishuli and Bhote Koshi Corridor Building inventory for the 26 August 2026 flash flood on the Nepal-China border. AOI: 1 km buffer around the Bhote Koshi and Trishuli river centrelines, 132.93 km² across Nuwakot and Rasuwa districts. Tasking Manager project 62904. Contents upperstream/buildings.geojson (+ .parquet): 13,663 footprints upperstream/building_density_h3_r8.geojson (+ .parquet): buildings per H3 res 8 cell (~0.7 km²)… See the full description on the dataset page: https://huggingface.co/datasets/hotosm/nepal_flood_2026.geospatial10K<n<100K0 likes126 downloads29d agoHugging Face06himalaya-ai /nepali-honorific-benchtabularn<1K0 likes99 downloads3mo agoHugging Face07ios-ioe /nepali-bias-dataset Nepali Bias Language Dataset Dataset Description A synthetic dataset of Nepali sentences labeled for bias categories including gender, religion, caste, regional, appearance, social status, political, age, and disability bias. Sentences were first labeled by LLMs (ChatGPT, Grok) prompted with real Nepali news context, then manually reviewed and corrected by human annotators. Dataset Summary Split Examples Train 1,362 Validation 292… See the full description on the dataset page: https://huggingface.co/datasets/ios-ioe/nepali-bias-dataset.tabulartext-classification1K<n<10K3 likes66 downloads4mo agoHugging Face08cloudfrm-site /nepali-tokenizer-corpustabular1M<n<10M0 likes60 downloads18d agoHugging Face09lilgoose7777 /aibharat_nepali_dataset_finalgatedtabular100K<n<1M0 likes56 downloads19d agoHugging Face10Aananda-giri /gorkhapatra-nepali-epaper Gorkhapatra Nepali E-Paper Corpus Per-article text extracted from PDF e-papers published on epaper.gorkhapatraonline.com, covering 11 newspaper slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal, loksewa, saturday, yuwamunch, gorkhapatra-125, other). Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.tabulartext-generation100K<n<1M0 likes51 downloads2mo agoHugging Face11open-llm-leaderboard /universalml__NepaliGPT-2.0-detailsgated Dataset Card for Evaluation run of universalml/NepaliGPT-2.0 Dataset automatically created during the evaluation run of model universalml/NepaliGPT-2.0 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/universalml__NepaliGPT-2.0-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face12dineshkarki /nepali-textbooks-corpus Nepali Textbooks Corpus for Grades 1-12 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 5634 Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12] Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.tabulartext-generation1K<n<10K2 likes43 downloads1y agoHugging Face13dineshkarki /nepali-tokenizer-corpustabular100K<n<1M0 likes43 downloads6mo agoHugging Face14Titung /cc100-nepali-cleaned CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text from CC-100 suitable for transformer pretraining. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Created: 2026-04-02 Pipeline Unicode normalisation (NFC + ftfy) Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate) Language ID — fastText lid.176.bin, confidence ≥ 0.7 Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.tabulartext-generation1M<n<10M2 likes40 downloads6mo agoHugging Face15biraj-bhusal /rakshak-nepali-toxicity-final 🛡️ RakshakAI — Augmented Nepali Toxicity Dataset The full augmented training dataset used to train the RakshakAI toxicity detection models. Contains 4,716 samples expanded from the curated 1,574 sample dataset through back-translation augmentation via English and Hindi as intermediate languages. For the clean curated dataset only, see rakshak-all-data-combined. 📄 Paper: RakshakAI: Multi-Label Toxicity Detection for Low-Resource Nepali Social Media Content Why this… See the full description on the dataset page: https://huggingface.co/datasets/biraj-bhusal/rakshak-nepali-toxicity-final.tabular1K<n<10K0 likes39 downloads3mo agoHugging Face16saliltambe /gemma4-e2b-nepali-sft-pairs Nepali SFT pairs for Gemma 4 E2B 468 (English prompt -> Nepali answer) pairs, the exact training data behind saliltambe/gemma-4-E2B-it-nepali-lora. Published so the training notebook can skip a ~13 minute generation step and so anyone reproducing it evaluates on the same held-out split. Provenance Prompts: English conversation openers from OpenAssistant/oasst1 (Apache-2.0, human-written), filtered to role == "prompter", parent_id is None, lang == "en". Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.tabularn<1K0 likes36 downloads7d agoHugging Face17jangedoo /nepali-redditDataset containing posts and comments from 3 subreddits NepalSocial, technepal and nepalstock. tabular1M<n<10M0 likes34 downloads3mo agoHugging Face18milanakdj /nepali-oov-distilledgated Nepali OOV-distilled subset (854 h) An OOV-dense distillation of Premal-12/c9nepali-audio-dataset2 (used with the author's permission), shipped in four variants: the original single-voice audio, a CPU-augmented copy, and 244 h re-rendered onto 1,842 real human speakers with Seed-VC. For Nepali ASR and TTS work. Filter with the variant field -- see Composition below. If you came here for speaker diversity, you want variant == "vc". What this is The source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-oov-distilled.audioautomatic-speech-recognition10K<n<100K1 likes29 downloads11d agoHugging Face19sumanpaudel1997 /nepali-asr-benchmark Nepali ASR Benchmark Per-utterance reference, hypothesis, WER, and CER for the six released Nepali ASR checkpoints evaluated on three independent test sets. Released alongside the paper Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition. Contents Field Type Description utterance_id string stable identifier {test_set}-{index} reference string NFC-normalised gold transcription (Devanagari) hypothesis string… See the full description on the dataset page: https://huggingface.co/datasets/sumanpaudel1997/nepali-asr-benchmark.tabularautomatic-speech-recognition10K<n<100K0 likes28 downloads4mo agoHugging Face20Khanalnishan /50k_nepali_dataset_chatbottabular10K<n<100K0 likes27 downloads6mo agoHugging Face21asal10 /nepal-supremecourt-judgments-cleanedtabular10K<n<100K0 likes26 downloads2mo agoHugging Face22rishikeshgautam /small-newscorpus-for-nepali-ged-fullstop-removedtabular100K<n<1M0 likes25 downloads2y agoHugging Face23cair-nepal /ai-bias-research-landscape Dataset Card: AI Bias Research Landscape Dataset Summary This dataset contains 692 curated bibliographic records of peer-reviewed and preprint publications on artificial intelligence (AI) and algorithmic bias, published between 2012 and 2026. Each record includes publication metadata (paper title, DOI, authors, author regions, affiliations, publication year, and research domain), author ORCID identifiers, and OpenAlex-derived metadata, including OpenAlex IDs… See the full description on the dataset page: https://huggingface.co/datasets/cair-nepal/ai-bias-research-landscape.tabularn<1K0 likes25 downloads3mo agoHugging Face24electricsheepasia /asia-aid-flows-financial-tracking-private-sector-nepal2 Financial tracking of private sector contributions Nepal 2015 Publisher: OCHA HQ · Source: HDX · License: cc-by-igo · Updated: 2023-05-02 Abstract Information on the private sector cash and in-kind contributions to humanitarian relief efforts in Nepal earthquake. Each row in this dataset represents tabular records. Temporal coverage is indicated by the unnamed_11, unnamed_12 column(s). Geographic scope: NPL, NEPAL-EARTHQUAKE. Curated into ML-ready Parquet format by… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-aid-flows-financial-tracking-private-sector-nepal2.tabulartabular-classificationn<1K0 likes24 downloads5mo agoHugging Face25sabin1234 /NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET Nepali Devanagari SFT Dataset — Final Clean Release A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments. Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns. Dataset at a Glance Property Value Total rows 100,000 Total conversation messages 200,000 Human messages 100,000 GPT messages 100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.texttext-generation100K<n<1M0 likes24 downloads1mo agoHugging Face26dineshkarki /textbooks-qa-nepali Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbooks-qa-nepali") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of 2 messages: human and gpt… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbooks-qa-nepali.tabularquestion-answering1K<n<10K1 likes23 downloads1y agoHugging Face27open-llm-leaderboard /shivam9980__NEPALI-LLM-detailsgated Dataset Card for Evaluation run of shivam9980/NEPALI-LLM Dataset automatically created during the evaluation run of model shivam9980/NEPALI-LLM The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/shivam9980__NEPALI-LLM-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face28lilgoose7777 /aibharat_nepali_datasetgatedtabular1K<n<10K0 likes21 downloads19d agoHugging Face29rishikeshgautam /small-newscorpus-for-nepali-gedtabular100K<n<1M1 likes20 downloads2y agoHugging Face30Titung /nepali-celeb-faces Usage from datasets import load_dataset ds = load_dataset("Titung/nepali-celeb-faces") imageobject-detection1K<n<10K0 likes19 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.