CoolFace
Datasetpublic

myamjechal/en_my_myanmar-xnli_small

myXNLI Preprocessed Dataset (Small Version) Overview This dataset is a smaller version of the myXNLI corpus, which extends the XNLI dataset to include Myanmar (Burmese) language. The dataset has been preprocessed, split into training, validation, and test sets, and is specifically designed for natural language inference (NLI) and machine translation tasks. Training set: 50,000 sentence pairs Validation set: 2,490 sentence pairs Test set: 5,010 sentence… See the full description on the dataset page: https://huggingface.co/datasets/myamjechal/en_my_myanmar-xnli_small.

sourceHugging Facecc-by-2.0updated 2y agoView on Hugging Face
1likes6downloads
Dataset Card

myXNLI Preprocessed Dataset (Small Version)

Overview

This dataset is a smaller version of the myXNLI corpus, which extends the XNLI dataset to include Myanmar (Burmese) language. The dataset has been preprocessed, split into training, validation, and test sets, and is specifically designed for natural language inference (NLI) and machine translation tasks.

  • —Training set: 50,000 sentence pairs
  • —Validation set: 2,490 sentence pairs
  • —Test set: 5,010 sentence pairs

The data is translated from English to Myanmar, with NLI and Genre labels retained from the original XNLI and MultiNLI datasets.

Dataset Structure

SplitNumber of Samples
Train50,000
Val2,490
Test5,010

Data Fields

  • —en (English): The English sentence in the pair (from sentence1_en).
  • —my (Myanmar): The corresponding Myanmar sentence in the pair (from sentence1_my).

License

This dataset is released under the Creative Commons Attribution Non Commercial 2.0 (CC BY-NC 2.0) license.

For more details, refer to the license.

Citation

If you use this dataset, please cite the original myXNLI corpus: myXNLI