datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xnli-eu
Dataset Card for XNLIeu
XNLIeu is an extension of XNLI translated from English to Basque. It has been designed as a cross-lingual dataset for the Natural Language Inference task, a text-classification task that consists on classifying pairs of sentences, a premise and a hypothesis, according to their semantic relation out of three possible labels: entailment, contradiction and neutral.
Dataset Details
Dataset Description
XNLI is a popular Natural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/xnli-eu.XNLIxnli2.0_train_arabicxnli_va
Dataset Summary
This dataset is a professional translation into Valencian of the Cross-lingual Natural Language Inference XNLI dataset.
XNLI-va is a collection of 5.010 sentence pairs annotated with textual entailment.
The original dataset was restricted to only non-commercial research purposes under the Creative Commons Attribution Non-commercial 4.0 International Public License.
Dataset Structure
premise: a string feature.
hypothesis: a string feature.
label: a… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/xnli_va.xnli2.0_thaiParallel-XNLIvarBrief dataset description:
Native: the native partition of XNLIeu (Heredia et al., 2024) adapted into three Basque dialects.
Test: the test partition of XNLI adapted into three Basque dialects.
All_dialects _together is a train/dev/test split that includes both native and test instances, stratified according to dialects.
xnli_gl
Dataset Card for XNLI_gl
XNLI_gl is an extension of XNLI translated to Galician. It has been designed as a cross-lingual dataset for the Natural Language Inference task, a text-classification task that consists on classifying pairs of sentences, a premise and a hypothesis, according to their semantic relation out of three possible labels: entailment, contradiction and neutral.
Curated by: Proxecto Nós
Language(s) (NLP): Galician
XNLI is restricted to only non-commercial research… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/xnli_gl.XNLI_Vietnamese_tripletsCitation:
@InProceedings{conneau2018xnli,
author = {Conneau, Alexis
and Rinott, Ruty
and Lample, Guillaume
and Williams, Adina
and Bowman, Samuel R.
and Schwenk, Holger
and Stoyanov, Veselin},
title = {XNLI: Evaluating Cross-lingual Sentence Representations},
booktitle = {Proceedings of the 2018 Conference on Empirical Methods
in Natural Language Processing},
year =… See the full description on the dataset page: https://huggingface.co/datasets/haiFrHust/XNLI_Vietnamese_triplets.xnli-translated-khm-pairclassificationxnli-translated-zsm-pairclassificationxnli2.0_train_englishxnli-translated-lao-pairclassificationburmese-xnli-mya-pairclassificationxnli2.0_train_thaixnli2.0_train_turkishXNLIvar
XNLIvar: Basque and Spanish variation-inclusive NLI
This repository contains the data and code used in the paper Lost in Variation? Evaluating NLI Performance in Basque and Spanish Geographical Variants
This paper evaluates the capacity of current language technologies to understand Basque and Spanish language varieties. We use Natural Language Inferenc (NLI) as a pivot task and introduce a novel, manually-curated parallel dataset in Basque and Spanish and their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/XNLIvar.xnli2.0_train_bulgarianxnli2.0_train_spanishxnli2.0_train_frenchxnli2.0_train_chinesexnli2.0_assamesexnli2.0_train_greekxnli2.0_train_russianxnli2.0_train_kannadaxnli2.0_chinesexnli2.0_train_hindixnli2.0_train_swahilixnli2.0_train_urdulanguage: ["Urdu"]
xnli2.0_train_sanskritxnli2.0_gujrati
