myamjechal/en_my_myanmar-xnli_small
myXNLI Preprocessed Dataset (Small Version) Overview This dataset is a smaller version of the myXNLI corpus, which extends the XNLI dataset to include Myanmar (Burmese) language. The dataset has been preprocessed, split into training, validation, and test sets, and is specifically designed for natural language inference (NLI) and machine translation tasks. Training set: 50,000 sentence pairs Validation set: 2,490 sentence pairs Test set: 5,010 sentence… See the full description on the dataset page: https://huggingface.co/datasets/myamjechal/en_my_myanmar-xnli_small.
myXNLI Preprocessed Dataset (Small Version)
Overview
This dataset is a smaller version of the myXNLI corpus, which extends the XNLI dataset to include Myanmar (Burmese) language. The dataset has been preprocessed, split into training, validation, and test sets, and is specifically designed for natural language inference (NLI) and machine translation tasks.
- Training set: 50,000 sentence pairs
- Validation set: 2,490 sentence pairs
- Test set: 5,010 sentence pairs
The data is translated from English to Myanmar, with NLI and Genre labels retained from the original XNLI and MultiNLI datasets.
Dataset Structure
Data Fields
- en (English): The English sentence in the pair (from
sentence1_en). - my (Myanmar): The corresponding Myanmar sentence in the pair (from
sentence1_my).
License
This dataset is released under the Creative Commons Attribution Non Commercial 2.0 (CC BY-NC 2.0) license.
For more details, refer to the license.
Citation
If you use this dataset, please cite the original myXNLI corpus: myXNLI
