CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aakashMeghwar01 /sindhi-corpus-505m Sindhi Corpus 505M The largest open-source, deduplicated Sindhi language pretraining corpus. ~505 million tokens across 742K documents, covering news, literature, legal, religious, encyclopedic, and web-crawled Sindhi text. Built for training Sindhi language models, tokenizers, and NLP tools. Dataset Summary Stat Value Documents ~742,379 Tokens (estimated) ~505 million Language Sindhi (sd) — Arabic script Format Parquet (single text column)… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/sindhi-corpus-505m.texttext-generation100K<n<1M2 likes165 downloads3mo agoHugging Face02aakkaasshh /vaigai-dataset Vaigai Dataset aakkaasshh/vaigai-dataset is an Indic-focused multilingual text corpus aggregated for tokenizer training, built with the goal of reaching tokenizer quality on par with dedicated Indic tokenizers (e.g. Sarvam AI's). One Parquet file per language is stored at the root of this repo (e.g. bn.parquet), with no nested folders. Every time new data is fetched for a language, it is merged with that language's existing file and deduplicated on the text column, so re-running… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/vaigai-dataset.tabulartext-generation1M<n<10M0 likes139 downloads2mo agoHugging Face03aakashMeghwar01 /Sindhi-Intelligence-Core-SFT 🧠 Sindhi Intelligence Core SFT This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning. 📊 Dataset Summary This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT). 📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.texttext-generation100K<n<1M1 likes58 downloads7mo agoHugging Face04AAkhoram /Persian-Wikipedia-Corpus Overview This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library. Usage from datasets import load_dataset dataset = load_dataset("codersan/Persian-Wikipedia-Corpus") Persian-Wikipedia-Corpus A complete copy of Persian Wikimedia pages, The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AAkhoram/Persian-Wikipedia-Corpus.tabulartext-generation1M<n<10M0 likes13 downloads2mo agoHugging Face05Aakanksha26 /train{ "use_cases": [ { "use_case": "General Reporting", "details": "N/A", "process_flow": "N/A" }, { "use_case": "Voice CDR", "details": "System shall handle errors and store CDRs in the database.", "process_flow": [ "User uploads voice CDR files to the system.", "System processes the files, validates, and ingests data into the database.", "System logs any errors encountered during processing." ] }, {… See the full description on the dataset page: https://huggingface.co/datasets/Aakanksha26/train.textsummarizationn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.