HeyDunaX/tay-vietnamese-nmt
Tày-Vietnamese Parallel Dataset The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages. Dataset Statistics Number of sentence pairs: 20,600 Average sentence… See the full description on the dataset page: https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt.
Tày-Vietnamese Parallel Dataset
The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages.
Dataset Statistics
- Number of sentence pairs: 20,600
- Average sentence length (Tày): 7.50 tokens
- Average sentence length (Vietnamese): 5.46 tokens
Limitations
- Limited data size due to the low-resource nature of the Tày language
- Coverage is limited to everyday communication and usage topics
Contributions
All authors contributed equally to the construction, preprocessing, and validation of the dataset.
Citation
@dataset{tran2025tayvie,
title = {Tày-Vietnamese Parallel Dataset},
author = {Tran, Nhan Duy and Ngo, Anh Phuong and Truong, Dat Thanh},
year = {2025},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt}
}