CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thegauravgiri /nepali-news-dataset 🇳🇵 Nepali News Dataset & NLP Corpus The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours. Repository: thegauravgiri/nepali-news-dataset Total Articles: 15,000+ full-text articles and growing Update Frequency: Every 4 hours via automated GitHub Actions pipelines Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding License: MIT License ⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.texttext-classification10K<n<100K1 likes7.2k downloads3h agoHugging Face02aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads7d agoHugging Face03aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads7d agoHugging Face04IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes973 downloads1y agoHugging Face05saileshbro /nepali-cs-asr Nepali–English Code-Switched ASR A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary. v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.audioautomatic-speech-recognition10K<n<100K1 likes894 downloads2mo agoHugging Face06cloudfrm-site /nepali-corpus-compiletext10M<n<100M0 likes809 downloads17d agoHugging Face07aarajbhattarai /nepali-agri-gov-instruct Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-agri-gov-instruct.text-generation0 likes789 downloads3m agoHugging Face08aarajbhattarai /rejected-nepali-agri-gov-instruct Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-agri-gov-instruct.text-generation0 likes760 downloads3m agoHugging Face09aarajbhattarai /unjudged-nepali-agri-gov-instruct Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-agri-gov-instruct.text-generation0 likes649 downloads10m agoHugging Face10NyayaLM /Nepali_pretraning_Corpustext10M<n<100M1 likes616 downloads1mo agoHugging Face11Aananda-giri /nepali_llm_datasets Nepali LLM Datasets This repository contains two configurations of Nepali LLM datasets: Configurations 1. Scrapy Engine Description: Contains data collected using a web scraping engine. Files: [List any specific files or formats] 2. Nepberta Description: This dataset is derived from the Nepberta project and contains cleaned data specifically related to the project. The dataset contains **cleaned text chunks of size ~50 mb ** of all… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/nepali_llm_datasets.text1M<n<10M1 likes601 downloads1y agoHugging Face12aarajbhattarai /unjudged-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.text-generation0 likes531 downloads7d agoHugging Face13DipeshChaudhary /nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes524 downloads11mo agoHugging Face14jangedoo /nepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency. text1M<n<10M0 likes519 downloads3mo agoHugging Face15tonibirat /neBrahma-Nepali-Pretrain-Corpus neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean, verified Devanagari Nepali text, drawn from four diverse sources and processed through an eight-stage cleaning and quality pipeline. This corpus serves as the training data for neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.texttext-generation10M<n<100M0 likes510 downloads3mo agoHugging Face16Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes491 downloads1y agoHugging Face17cloudfrm-site /sangraha_nepalitext10M<n<100M0 likes481 downloads17d agoHugging Face18DipeshChaudhary /muril-nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes469 downloads11mo agoHugging Face19mridul3301 /nepali-text-corpus-64 Nepali Text Dataset Overview The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset Details Total Articles: ~6.4 million Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.text1M<n<10M5 likes454 downloads2y agoHugging Face20paudelapil /nepali_asr_dataaudio1K<n<10K0 likes447 downloads2mo agoHugging Face21himalaya-ai /nepalipixel-synthetic-ocr-benchmark NepaliPixel Benchmark Dataset Model Card Overview The output_benchmark directory contains a synthetic OCR benchmark dataset for Nepali (Devanagari) script generated using the Nepali Pixel pipeline. The dataset is designed for evaluation of OCR models with minimal augmentation noise and fully rendered pages. Data Samples: Approximately 15,000 image‑text pairs (generated with -n 15000). Granularity: Includes all five levels – word, sentence… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepalipixel-synthetic-ocr-benchmark.imageimage-to-text10K<n<100K5 likes431 downloads3mo agoHugging Face22chintalaswathi /nepali-audio-deepfake-datasetaudio1K<n<10K0 likes405 downloads2mo agoHugging Face23himalaya-ai /nepali-corpus-compiletext10M<n<100M3 likes395 downloads6mo agoHugging Face24himalaya-ai /nepali-roman-pretraintext10M<n<100M0 likes390 downloads6mo agoHugging Face25Firoj112 /nepali-kokoro-ft-data Nepali Kokoro Fine-Tuning Dataset This is a sharded, processed dataset containing Nepali voice data for Kokoro TTS fine-tuning. audio10K<n<100K0 likes379 downloads2mo agoHugging Face26himalaya-ai /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.texttext-generation1K<n<10K4 likes362 downloads26d agoHugging Face27NepaliAI /Nepali-HealthChattextquestion-answering10K<n<100K3 likes360 downloads3y agoHugging Face28rdabin /nepali_dataset_llmtext10M<n<100M0 likes344 downloads1y agoHugging Face29milanakdj /nepali-audio-reserve-r6gated Nepali two-speaker conversation chunks ~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.audioautomatic-speech-recognition10K<n<100K0 likes338 downloads13d agoHugging Face30aarajbhattarai /nepali-fruit-rerun Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.texttext-generation1K<n<10K0 likes323 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.