CoolFace
Datasetpublic

HeyDunaX/tay-vietnamese-nmt

Tày-Vietnamese Parallel Dataset The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages. Dataset Statistics Number of sentence pairs: 20,600 Average sentence… See the full description on the dataset page: https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt.

sourceHugging Facecc-by-nc-4.0updated 6mo agoView on Hugging Face
1likes33downloads
Dataset Card

Tày-Vietnamese Parallel Dataset

The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages.

Dataset Statistics

  • —Number of sentence pairs: 20,600
  • —Average sentence length (Tày): 7.50 tokens
  • —Average sentence length (Vietnamese): 5.46 tokens

Limitations

  • —Limited data size due to the low-resource nature of the Tày language
  • —Coverage is limited to everyday communication and usage topics

Contributions

All authors contributed equally to the construction, preprocessing, and validation of the dataset.

Citation

bibtex
@dataset{tran2025tayvie,
  title     = {Tày-Vietnamese Parallel Dataset},
  author    = {Tran, Nhan Duy and Ngo, Anh Phuong and Truong, Dat Thanh},
  year      = {2025},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt}
}