datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sindhi-corpus-505m
Sindhi Corpus 505M
The largest open-source, deduplicated Sindhi language pretraining corpus.
~505 million tokens across 742K documents, covering news, literature, legal, religious, encyclopedic, and web-crawled Sindhi text. Built for training Sindhi language models, tokenizers, and NLP tools.
Dataset Summary
Stat
Value
Documents
~742,379
Tokens (estimated)
~505 million
Language
Sindhi (sd) — Arabic script
Format
Parquet (single text column)… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/sindhi-corpus-505m.vaigai-dataset
Vaigai Dataset
aakkaasshh/vaigai-dataset is an Indic-focused multilingual text corpus aggregated for tokenizer training, built with the goal of reaching tokenizer quality on par with dedicated Indic tokenizers (e.g. Sarvam AI's).
One Parquet file per language is stored at the root of this repo (e.g. bn.parquet), with no nested folders.
Every time new data is fetched for a language, it is merged with that language's existing file and deduplicated on the text column, so re-running… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/vaigai-dataset.Sindhi-Intelligence-Core-SFT
🧠 Sindhi Intelligence Core SFT
This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning.
📊 Dataset Summary
This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT).
📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.Persian-Wikipedia-Corpus
Overview
This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library.
Usage
from datasets import load_dataset
dataset = load_dataset("codersan/Persian-Wikipedia-Corpus")
Persian-Wikipedia-Corpus
A complete copy of Persian Wikimedia pages, The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AAkhoram/Persian-Wikipedia-Corpus.train{
"use_cases": [
{
"use_case": "General Reporting",
"details": "N/A",
"process_flow": "N/A"
},
{
"use_case": "Voice CDR",
"details": "System shall handle errors and store CDRs in the database.",
"process_flow": [
"User uploads voice CDR files to the system.",
"System processes the files, validates, and ingests data into the database.",
"System logs any errors encountered during processing."
]
},
{… See the full description on the dataset page: https://huggingface.co/datasets/Aakanksha26/train.
