CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thegauravgiri /nepali-news-dataset 🇳🇵 Nepali News Dataset & NLP Corpus The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours. Repository: thegauravgiri/nepali-news-dataset Total Articles: 15,000+ full-text articles and growing Update Frequency: Every 4 hours via automated GitHub Actions pipelines Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding License: MIT License ⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.texttext-classification10K<n<100K1 likes6.1k downloads8h agoHugging Face02nepfaff /scenesmith-example-scenes SceneSmith Example Scenes Project Page | Paper | Code Example scenes generated by SceneSmith, a hierarchical agentic framework for constructing simulation-ready indoor environments from natural language prompts. This dataset contains all scenes from the SceneSmith method (and its ablations) used in the paper evaluations. Each scene is a complete simulation-ready environment with 3D assets (including VLM-estimated physical properties), collision meshes, floor plans, and scene… See the full description on the dataset page: https://huggingface.co/datasets/nepfaff/scenesmith-example-scenes.robotics1K<n<10K11 likes2.5k downloads4mo agoHugging Face03aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads4d agoHugging Face04aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads4d agoHugging Face05Neph0s /CoSER CoSER Dataset Overview CoSER is a high-quality dataset for role-playing LLMs, sourced from 771 renowned novels. The dataset contains authentic multi-turn, multi-character dialogues extracted from acclaimed literary works. Key Features Authentic Content: Unlike synthetic datasets, CoSER extracts real dialogues from literature, maintaining high fidelity to the original works. The dialogues are inherently multi-turn and multi-character, exhibiting natural… See the full description on the dataset page: https://huggingface.co/datasets/Neph0s/CoSER.86 likes1.4k downloads1y agoHugging Face06IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes990 downloads1y agoHugging Face07Neptune615 /Wild-City WildCity Dataset WildCity is a real-world city-scale multimodal dataset for street-view reconstruction, simulation, and spatial intelligence. It is collected from autonomous-driving fleet logs across multiple U.S. cities and contains surround-view RGB images, LiDAR, calibration, ego and sensor poses, object annotations, semantic masks, and processed reconstruction assets. This repository hosts the initial public release of WildCity. This version does not include the full raw… See the full description on the dataset page: https://huggingface.co/datasets/Neptune615/Wild-City.image-to-3d4 likes976 downloads3mo agoHugging Face08saileshbro /nepali-cs-asr Nepali–English Code-Switched ASR A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary. v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.audioautomatic-speech-recognition10K<n<100K1 likes884 downloads2mo agoHugging Face09cloudfrm-site /nepali-corpus-compiletext10M<n<100M0 likes796 downloads15d agoHugging Face10paudelapil /nepali_asr_dataaudio1K<n<10K0 likes624 downloads2mo agoHugging Face11Aananda-giri /nepali_llm_datasets Nepali LLM Datasets This repository contains two configurations of Nepali LLM datasets: Configurations 1. Scrapy Engine Description: Contains data collected using a web scraping engine. Files: [List any specific files or formats] 2. Nepberta Description: This dataset is derived from the Nepberta project and contains cleaned data specifically related to the project. The dataset contains **cleaned text chunks of size ~50 mb ** of all… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/nepali_llm_datasets.text1M<n<10M1 likes607 downloads1y agoHugging Face12DipeshChaudhary /nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes596 downloads11mo agoHugging Face13NyayaLM /Nepali_pretraning_Corpustext10M<n<100M1 likes596 downloads28d agoHugging Face14damo-da /oag-nepal-audit-reports OAG Nepal Audit Reports — Nepali transcripts and ruled tables Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements. The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.documenttext-retrieval10M<n<100M0 likes554 downloads20d agoHugging Face15aarajbhattarai /unjudged-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.text-generation0 likes531 downloads4d agoHugging Face16PNNL /NEPATEC3.0gated National Environmental Policy Act Text Corpus (NEPATEC3.0) Dataset Description The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for… See the full description on the dataset page: https://huggingface.co/datasets/PNNL/NEPATEC3.0.text-generation100K<n<1M9 likes529 downloads14d agoHugging Face17Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes479 downloads1y agoHugging Face18DipeshChaudhary /muril-nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes473 downloads11mo agoHugging Face19cloudfrm-site /sangraha_nepalitext10M<n<100M0 likes471 downloads15d agoHugging Face20tonibirat /neBrahma-Nepali-Pretrain-Corpus neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean, verified Devanagari Nepali text, drawn from four diverse sources and processed through an eight-stage cleaning and quality pipeline. This corpus serves as the training data for neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.texttext-generation10M<n<100M0 likes466 downloads3mo agoHugging Face21mridul3301 /nepali-text-corpus-64 Nepali Text Dataset Overview The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset Details Total Articles: ~6.4 million Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.text1M<n<10M5 likes455 downloads2y agoHugging Face22jangedoo /nepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency. text1M<n<10M0 likes455 downloads3mo agoHugging Face23himalaya-ai /nepalipixel-synthetic-ocr-benchmark NepaliPixel Benchmark Dataset Model Card Overview The output_benchmark directory contains a synthetic OCR benchmark dataset for Nepali (Devanagari) script generated using the Nepali Pixel pipeline. The dataset is designed for evaluation of OCR models with minimal augmentation noise and fully rendered pages. Data Samples: Approximately 15,000 image‑text pairs (generated with -n 15000). Granularity: Includes all five levels – word, sentence… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepalipixel-synthetic-ocr-benchmark.imageimage-to-text10K<n<100K5 likes429 downloads3mo agoHugging Face24chintalaswathi /nepali-audio-deepfake-datasetaudio1K<n<10K0 likes396 downloads2mo agoHugging Face25himalaya-ai /nepali-corpus-compiletext10M<n<100M3 likes391 downloads6mo agoHugging Face26nepfaff /scenesmith-preprocessed-data SceneSmith Preprocessed Data Preprocessed 3D assets for use with SceneSmith, a VLM-agent-based system for generating physically realistic, interactive indoor scenes. ArtVIP (Articulated Objects) Simulation-ready articulated objects (cabinets, drawers, appliances, etc.) converted from the ArtVIP dataset. Assets have been converted from USD to Drake SDFormat using mesh-to-sim-asset, with: Drake SDFormat (.sdf) model files with articulated joints Visual meshes in… See the full description on the dataset page: https://huggingface.co/datasets/nepfaff/scenesmith-preprocessed-data.1 likes389 downloads4mo agoHugging Face27himalaya-ai /nepali-roman-pretraintext10M<n<100M0 likes388 downloads6mo agoHugging Face28AnkitSanjyal /nepse-market-data NEPSE Market Data Daily and tick-level data from the Nepal Stock Exchange, captured by an open-source pipeline and validated at every layer boundary. Daily prices reach back to 1995-07-20. Tick data starts 2026-08-19. That gap is a property of the source, not a backlog - see Coverage below. Tables Path Grain Source Coverage eod_history_ext/ symbol x day ShareSansar (second source) 1995-07-20 -> present, 653 symbols eod_price/ symbol x day NEPSE… See the full description on the dataset page: https://huggingface.co/datasets/AnkitSanjyal/nepse-market-data.tabular100K<n<1M0 likes384 downloads28d agoHugging Face29himalaya-ai /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/hermes-function-calling-nepali.texttext-generation1K<n<10K4 likes357 downloads23d agoHugging Face30NepaliAI /Nepali-HealthChattextquestion-answering10K<n<100K3 likes351 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.