datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kashmiri-language-text-datasetThis is a Dataset for Kashmiri language containing Kashmiri words their phonemes , along with their meaning in ENGlISH and HINDI
and an example sentence in english . The phoenmes are written phonemes present in wordphonemes-meaning.csv
Kashmiri__Text_Corpus_Dataset
Comprehensive Kashmiri Text Analysis Report
1. File Information
Filename:FULL CORPUS KASHMIRI TEXT.txt
Size: 15 Mb (15066.21 KB) (15,427,794 bytes)
Created: 2024-10-29 19:35:41
Modified: 2024-10-29 19:35:41
encoding: utf-8
language_hint: kashmiri
2. Text Statistics
Total Characters: 8,560,982
Total Words: 1,691,085 (1.6 million)
Total Lines: 22,882 (one line collection of many sentences)
Unique Words: 57,073
Vocabulary Density: 3.37%
Words per Line:… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri__Text_Corpus_Dataset.3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM
DATASET NAME: KS-LIT-3M
Kashmiri Pretraining Dataset
This repository hosts a meticulously processed Kashmiri text dataset, specifically designed for pretraining Large Language Models (LLMs) from scratch. The dataset has undergone extensive cleaning and preprocessing to ensure high quality and suitability for robust model training.
Dataset Description
This dataset consists of a continuous one stream of Kashmiri text, cleaned to remove English… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM.KS-PRET-5M_5_million_kashmiri_Pretrainning_LLM_dataset_12M_tokens_2026
KS-PRET-5M: Kashmiri Pretraining Corpus
5,090,244 words · ~12,130,000 subword tokens · 295,433 vocabulary · April 2026
Building the complete AI stack for Kashmiri — 7 million speakers, virtually no prior computational resources.
Dataset Summary
KS-PRET-5M is a large-scale, deeply cleaned Kashmiri language pretraining corpus — the largest publicly available dataset for the Kashmiri language. All text is formatted as a single continuous stream, the standard… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/KS-PRET-5M_5_million_kashmiri_Pretrainning_LLM_dataset_12M_tokens_2026.Kashmiri_Text_Corpus_Cleaned_2025_HNM
Overview
A comprehensive linear text corpus of the Kashmiri language, optimized for large language model (LLM) pre-training.
Corpus Text Analysis Report
The corpus contains approximately 2 million words, with over 91,000 unique words
The entire corpus is formatted as a single continuous line of text, ideal for LLM Pre-training
The cleaning process removed about 4.4% of characters while preserving 99.3% of words
Vocabulary richness… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM.Kashmiri_Poet_dataset_Shruk
Shruks Analysis
Total Number of Shruks:
1552
Total Lines
8283
Total Number of Words:
35992
Total Number of Characters:
195236
2. Dataset Structure for NLP Tasks
You can use the Shruks dataset for various NLP tasks.
You can apply this methodology to any task, whether it's text classification, sentiment analysis, or machine translation.
4. Potential Use Cases for This Dataset
Text Classification: Classifying Shruks… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Poet_dataset_Shruk.
