CoolFace
Datasetpublic

HeyDunaX/tay-vietnamese-nmt

Tày-Vietnamese Parallel Dataset The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages. Dataset Statistics Number of sentence pairs: 20,600 Average sentence… See the full description on the dataset page: https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt.

sourceHugging Facecc-by-nc-4.0updated 6mo agoView on Hugging Face
1likes33downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
HeyDunaX/tay-vietnamese-nmt · CoolFace