CoolFace
Datasetpublicgated

TajikNLPWorld/TajPersParallelCorpus

Dataset Card for Tajik–Persian Parallel Corpus Dataset Details Dataset Description The Tajik–Persian Parallel Corpus is a large-scale parallel corpus containing 328,253 aligned Tajik–Persian sentence pairs collected from multiple sources, including news, poetry, prose, lexical resources, and named-entity lists. It is intended for machine translation, cross-lingual retrieval, linguistic analysis, tokenizer evaluation, and other NLP tasks. Curated… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/TajPersParallelCorpus.

sourceHugging Faceotherupdated 29d agoView on Hugging Face
0likes27downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

TajikNLPWorld/TajPersParallelCorpus · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.