nep
Datasets
All datasets matching “nep”nepali-news-dataset
🇳🇵 Nepali News Dataset & NLP Corpus
The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours.
Repository: thegauravgiri/nepali-news-dataset
Total Articles: 15,000+ full-text articles and growing
Update Frequency: Every 4 hours via automated GitHub Actions pipelines
Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding
License: MIT License
⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.scenesmith-example-scenes
SceneSmith Example Scenes
Project Page | Paper | Code
Example scenes generated by SceneSmith, a hierarchical agentic framework for constructing simulation-ready indoor environments from natural language prompts.
This dataset contains all scenes from the SceneSmith method (and its ablations) used in the paper evaluations. Each scene is a complete simulation-ready environment with 3D assets (including VLM-estimated physical properties), collision meshes, floor plans, and scene… See the full description on the dataset page: https://huggingface.co/datasets/nepfaff/scenesmith-example-scenes.nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.CoSER
CoSER Dataset
Overview
CoSER is a high-quality dataset for role-playing LLMs, sourced from 771 renowned novels. The dataset contains authentic multi-turn, multi-character dialogues extracted from acclaimed literary works.
Key Features
Authentic Content: Unlike synthetic datasets, CoSER extracts real dialogues from literature, maintaining high fidelity to the original works. The dialogues are inherently multi-turn and multi-character, exhibiting natural… See the full description on the dataset page: https://huggingface.co/datasets/Neph0s/CoSER.Nepali-Text-Corpus
Nepali Text Corpus
Overview
Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a
diverse range of text types, including news articles, blogs, and more, making it an invaluable
resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP)
and computational linguistics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.
