CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads7d agoHugging Face02aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads7d agoHugging Face03IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes973 downloads1y agoHugging Face04saileshbro /nepali-cs-asr Nepali–English Code-Switched ASR A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary. v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.audioautomatic-speech-recognition10K<n<100K1 likes894 downloads2mo agoHugging Face05cloudfrm-site /nepali-corpus-compiletext10M<n<100M0 likes809 downloads18d agoHugging Face06NyayaLM /Nepali_pretraning_Corpustext10M<n<100M1 likes616 downloads1mo agoHugging Face07Aananda-giri /nepali_llm_datasets Nepali LLM Datasets This repository contains two configurations of Nepali LLM datasets: Configurations 1. Scrapy Engine Description: Contains data collected using a web scraping engine. Files: [List any specific files or formats] 2. Nepberta Description: This dataset is derived from the Nepberta project and contains cleaned data specifically related to the project. The dataset contains **cleaned text chunks of size ~50 mb ** of all… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/nepali_llm_datasets.text1M<n<10M1 likes601 downloads1y agoHugging Face08DipeshChaudhary /nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes524 downloads11mo agoHugging Face09jangedoo /nepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency. text1M<n<10M0 likes519 downloads3mo agoHugging Face10tonibirat /neBrahma-Nepali-Pretrain-Corpus neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean, verified Devanagari Nepali text, drawn from four diverse sources and processed through an eight-stage cleaning and quality pipeline. This corpus serves as the training data for neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.texttext-generation10M<n<100M0 likes510 downloads3mo agoHugging Face11Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes491 downloads1y agoHugging Face12cloudfrm-site /sangraha_nepalitext10M<n<100M0 likes481 downloads18d agoHugging Face13DipeshChaudhary /muril-nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes469 downloads11mo agoHugging Face14mridul3301 /nepali-text-corpus-64 Nepali Text Dataset Overview The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset Details Total Articles: ~6.4 million Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.text1M<n<10M5 likes454 downloads2y agoHugging Face15himalaya-ai /nepalipixel-synthetic-ocr-benchmark NepaliPixel Benchmark Dataset Model Card Overview The output_benchmark directory contains a synthetic OCR benchmark dataset for Nepali (Devanagari) script generated using the Nepali Pixel pipeline. The dataset is designed for evaluation of OCR models with minimal augmentation noise and fully rendered pages. Data Samples: Approximately 15,000 image‑text pairs (generated with -n 15000). Granularity: Includes all five levels – word, sentence… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepalipixel-synthetic-ocr-benchmark.imageimage-to-text10K<n<100K5 likes431 downloads3mo agoHugging Face16himalaya-ai /nepali-corpus-compiletext10M<n<100M3 likes395 downloads6mo agoHugging Face17himalaya-ai /nepali-roman-pretraintext10M<n<100M0 likes390 downloads6mo agoHugging Face18Firoj112 /nepali-kokoro-ft-data Nepali Kokoro Fine-Tuning Dataset This is a sharded, processed dataset containing Nepali voice data for Kokoro TTS fine-tuning. audio10K<n<100K0 likes379 downloads2mo agoHugging Face19himalaya-ai /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.texttext-generation1K<n<10K4 likes362 downloads26d agoHugging Face20NepaliAI /Nepali-HealthChattextquestion-answering10K<n<100K3 likes360 downloads3y agoHugging Face21rdabin /nepali_dataset_llmtext10M<n<100M0 likes344 downloads1y agoHugging Face22milanakdj /nepali-audio-reserve-r6gated Nepali two-speaker conversation chunks ~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.audioautomatic-speech-recognition10K<n<100K0 likes338 downloads13d agoHugging Face23aarajbhattarai /nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.texttext-generation1K<n<10K0 likes323 downloads4d agoHugging Face24aarajbhattarai /rejected-nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.texttext-generation1K<n<10K0 likes321 downloads4d agoHugging Face25himalaya-ai /nepali-pretrain-corpustext10M<n<100M3 likes284 downloads6mo agoHugging Face26himalaya-ai /cc100-nepali CC-100 Nepali — Cleaned & Deduplicated Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from CC-100, suitable for transformer pretraining. Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned. Statistics Split Sentences train 4,736,157 validation 48,328 test 48,329 total 4,832,814 Token Statistics (train split) Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.tabulartext-generation1M<n<10M0 likes267 downloads6mo agoHugging Face27himalaya-ai /sangraha_nepalitext10M<n<100M0 likes238 downloads6mo agoHugging Face28Firoj112 /nepali-asr-whisper Nepali ASR Dataset (FLAC, Prepared for Whisper Fine-Tuning) This dataset is a preprocessed and ready-to-use version of the OpenSLR Nepali Automatic Speech Recognition (ASR) corpus, repackaged and standardized to facilitate Whisper model fine-tuning and other speech-to-text experiments. It provides audio-text pairs in Nepali language, with all audio stored as .flac files and transcriptions in Devanagari script. The dataset follows the Hugging Face datasets format for seamless use… See the full description on the dataset page: https://huggingface.co/datasets/Firoj112/nepali-asr-whisper.audio1M<n<10M1 likes231 downloads11mo agoHugging Face29cloudfrm-site /nepali-pretrain-corpustext10M<n<100M0 likes222 downloads18d agoHugging Face30mteb /NepaliNewsClassification NepaliNewsClassification An MTEB dataset Massive Text Embedding Benchmark A Nepali dataset for 7500 news articles Task category t2c Domains News, Written Reference https://github.com/goru001/nlp-for-nepali How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["NepaliNewsClassification"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NepaliNewsClassification.texttext-classification1K<n<10K0 likes221 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.