CoolFace
Datasetpublic

arnizamani/Sindhi-texts-big-dataset

Sindhi Texts (big dataset) A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia, newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora. 3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5 characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
4likes9kdownloads

arnizamani/Sindhi-texts-big-dataset · main · files are served by the source, never re-hosted here

arnizamani/Sindhi-texts-big-dataset · CoolFace