CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01injilashah /Kashmiri-language-text-datasetThis is a Dataset for Kashmiri language containing Kashmiri words their phonemes , along with their meaning in ENGlISH and HINDI and an example sentence in english . The phoenmes are written phonemes present in wordphonemes-meaning.csv translation0 likes401 downloads2y agoHugging Face02Omarrran /Kashmiri__Text_Corpus_Datasetgated Comprehensive Kashmiri Text Analysis Report 1. File Information Filename:FULL CORPUS KASHMIRI TEXT.txt Size: 15 Mb (15066.21 KB) (15,427,794 bytes) Created: 2024-10-29 19:35:41 Modified: 2024-10-29 19:35:41 encoding: utf-8 language_hint: kashmiri 2. Text Statistics Total Characters: 8,560,982 Total Words: 1,691,085 (1.6 million) Total Lines: 22,882 (one line collection of many sentences) Unique Words: 57,073 Vocabulary Density: 3.37% Words per Line:… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri__Text_Corpus_Dataset.texttext-generation10K<n<100K3 likes14 downloads2y agoHugging Face03Omarrran /3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNMgated DATASET NAME: KS-LIT-3M Kashmiri Pretraining Dataset This repository hosts a meticulously processed Kashmiri text dataset, specifically designed for pretraining Large Language Models (LLMs) from scratch. The dataset has undergone extensive cleaning and preprocessing to ensure high quality and suitability for robust model training. Dataset Description This dataset consists of a continuous one stream of Kashmiri text, cleaned to remove English… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM.texttext-generationn<1K4 likes13 downloads5mo agoHugging Face04Omarrran /KS-PRET-5M_5_million_kashmiri_Pretrainning_LLM_dataset_12M_tokens_2026gated KS-PRET-5M: Kashmiri Pretraining Corpus 5,090,244 words · ~12,130,000 subword tokens · 295,433 vocabulary · April 2026 Building the complete AI stack for Kashmiri — 7 million speakers, virtually no prior computational resources. Dataset Summary KS-PRET-5M is a large-scale, deeply cleaned Kashmiri language pretraining corpus — the largest publicly available dataset for the Kashmiri language. All text is formatted as a single continuous stream, the standard… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/KS-PRET-5M_5_million_kashmiri_Pretrainning_LLM_dataset_12M_tokens_2026.texttext-generationn<1K1 likes10 downloads5mo agoHugging Face05Omarrran /Kashmiri_Text_Corpus_Cleaned_2025_HNMgated Overview A comprehensive linear text corpus of the Kashmiri language, optimized for large language model (LLM) pre-training. Corpus Text Analysis Report The corpus contains approximately 2 million words, with over 91,000 unique words The entire corpus is formatted as a single continuous line of text, ideal for LLM Pre-training The cleaning process removed about 4.4% of characters while preserving 99.3% of words Vocabulary richness… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM.texttext-generationn<1K0 likes7 downloads9mo agoHugging Face06Omarrran /Kashmiri_Poet_dataset_Shrukgated Shruks Analysis Total Number of Shruks: 1552 Total Lines 8283 Total Number of Words: 35992 Total Number of Characters: 195236 2. Dataset Structure for NLP Tasks You can use the Shruks dataset for various NLP tasks. You can apply this methodology to any task, whether it's text classification, sentiment analysis, or machine translation. 4. Potential Use Cases for This Dataset Text Classification: Classifying Shruks… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Poet_dataset_Shruk.texttext-generation1K<n<10K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.