datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
snli-training-paraphrase-augmentation
SNLI Training Paraphrase Augmentation
Purpose
This dataset contains new paraphrases created for training paraphrase augmentation during Phase B of the research project.
It was not used as an NLI evaluation set.
It was not used for paraphrase consistency evaluation.
It is separate from the published SNLI Paraphrase Bank used for evaluation. The generation records report zero collisions with that evaluation bank.
The CSV contains only the new augmentation rows. It… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/snli-training-paraphrase-augmentation.snli-paraphrase-bank
SNLI Paraphrase Bank
Overview
This repository contains paraphrases for hypotheses in the SNLI training split.
Each original hypothesis has up to five paraphrases. Every retained paraphrase passed two automated checks. The first check tested semantic equivalence with the original hypothesis. The second check tested whether the relation between the premise and the paraphrase matched the original gold label.
The dataset contains 2,552,844 paraphrase rows from 533,653… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/snli-paraphrase-bank.snli-en-sk-nllb-25000
