CoolFace
Datasetpublic

arnizamani/Sindhi-texts-big-dataset

Sindhi Texts (big dataset) A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia, newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora. 3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5 characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
4likes9kdownloads
settings

This repository belongs to arnizamani on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSindhi-texts-big-dataset
visibilitypublic
licencemit
gatedno
ownerarnizamani
Account settings
arnizamani/Sindhi-texts-big-dataset · CoolFace