CoolFace
Datasetpublicgated

Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM

Overview A comprehensive linear text corpus of the Kashmiri language, optimized for large language model (LLM) pre-training. Corpus Text Analysis Report The corpus contains approximately 2 million words, with over 91,000 unique words The entire corpus is formatted as a single continuous line of text, ideal for LLM Pre-training The cleaning process removed about 4.4% of characters while preserving 99.3% of words… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM.

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes7downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.