CoolFace
Datasetpublic

HeyDunaX/tay-vietnamese-nmt

Tày-Vietnamese Parallel Dataset The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages. Dataset Statistics Number of sentence pairs: 20,600 Average sentence… See the full description on the dataset page: https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt.

sourceHugging Facecc-by-nc-4.0updated 6mo agoView on Hugging Face
1likes33downloads
5 commits on main
67974506mo ago

Update README.md

HeyDunaX
2b04e139mo ago

Update README.md

HeyDunaX
52299709mo ago

Create README.md

HeyDunaX
91ff7049mo ago

Upload dataset

HeyDunaX
21d6cb69mo ago

initial commit

HeyDunaX