CoolFace
Datasetpublicgated

TajikNLPWorld/tajik-wiki-corpus

Dataset Card for Tajik Wikipedia Corpus Dataset Details Dataset Description The Tajik Wikipedia Corpus is a collection of 79,985 Wikipedia articles in the Tajik language, totaling approximately 101.6 million characters and 15.3 million words. The data has been extracted from the Tajik Wikipedia dump and processed to ensure clean, well‑structured text suitable for NLP applications. The corpus includes articles with titles, categories, and metadata.… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-wiki-corpus.

sourceHugging Facecc-by-sa-4.0updated 1mo agoView on Hugging Face
1likes35downloads

TajikNLPWorld/tajik-wiki-corpus · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.