CoolFace
Datasetpublicgated

TajikNLPWorld/tajik-wiki-corpus

Dataset Card for Tajik Wikipedia Corpus Dataset Details Dataset Description The Tajik Wikipedia Corpus is a collection of 79,985 Wikipedia articles in the Tajik language, totaling approximately 101.6 million characters and 15.3 million words. The data has been extracted from the Tajik Wikipedia dump and processed to ensure clean, well‑structured text suitable for NLP applications. The corpus includes articles with titles, categories, and metadata.… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-wiki-corpus.

sourceHugging Facecc-by-sa-4.0updated 29d agoView on Hugging Face
1likes35downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.