mridul3301/nepali-text-corpus-64
Nepali Text Dataset Overview The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset Details Total Articles: ~6.4 million… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.
Nepali Text Dataset
Overview
The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics.
Dataset Details
- Total Articles: ~6.4 million
- Language: Nepali
- Size: 27.5 GB (in csv)
- Source: Collected from various Nepali news websites, blogs, and other online platforms.
Features
- Diversity: Includes a wide range of topics such as politics, culture, technology, and entertainment.
- Rich Vocabulary: Captures the nuances of the Nepali language, including idiomatic expressions and regional dialects.
- Clean: Is clean and work ready.
Purpose
Primary purpose of generating this dataset is : *Language modeling*.
