CoolFace
Datasetpublic

tonibirat/neBrahma-Nepali-Pretrain-Corpus

neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean, verified Devanagari Nepali text, drawn from four diverse sources and processed through an eight-stage cleaning and quality pipeline. This corpus serves as the training data for neBrahma-llm - a… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/neBrahma-Nepali-Pretrain-Corpus.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes522downloads

tonibirat/neBrahma-Nepali-Pretrain-Corpus · main · files are served by the source, never re-hosted here