TajikNLPWorld/tajik-wiki-corpus
Dataset Card for Tajik Wikipedia Corpus Dataset Details Dataset Description The Tajik Wikipedia Corpus is a collection of 79,985 Wikipedia articles in the Tajik language, totaling approximately 101.6 million characters and 15.3 million words. The data has been extracted from the Tajik Wikipedia dump and processed to ensure clean, well‑structured text suitable for NLP applications. The corpus includes articles with titles, categories, and metadata.… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-wiki-corpus.
This repository belongs to TajikNLPWorld on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
