CoolFace
Datasetpublic

Sakonii/nepalitext-language-model-dataset

Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.

sourceHugging Facecc0-1.0updated 1y agoView on Hugging Face
8likes502downloads

Sakonii/nepalitext-language-model-dataset · main · files are served by the source, never re-hosted here