datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ordering_Constrained_ParaphrasesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Ordering_Constrained_Paraphrases.Reorient_Block_ParaphrasesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Reorient_Block_Paraphrases.harvest_apples_with_agilex_piper_sim_ee_paraphrases20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"observation.state": {
"dtype": "float32",
"fps": 25,
"shape": [
8
],
"names": [
"ee.x",
"ee.y",
"ee.z",
"ee.roll",
"ee.pitch",
"ee.yaw"… See the full description on the dataset page: https://huggingface.co/datasets/Faless/harvest_apples_with_agilex_piper_sim_ee_paraphrases20.autoencoder-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models.
It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses).
The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.paraphraseger-backtrans-paraphrase
German Backtranslated Paraphrase Dataset
This is a dataset of more than 21 million German paraphrases.
These are text pairs that have the same meaning but are expressed with different words.
The source of the paraphrases are different parallel German / English text corpora.
The English texts were machine translated back into German to obtain the paraphrases.
This dataset can be used for example to train semantic text embeddings.
To do this, for example, SentenceTransformers
and the… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/ger-backtrans-paraphrase.medical_questions_paraphrasesopen-australian-legal-qa-paraphrased-easy-geminicluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details
Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid
Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details.open-australian-legal-qa-paraphrased-hard-gemini-with-embopen-australian-legal-qa-paraphrased-easy-gemini-with-embopen-australian-legal-qa-paraphrased-hard-geminibangla-health-related-paraphrased-dataset
Dataset Card for "BanglaHealthParaphrase"
BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.slop-paraphrase-pairs-v2books3_basic_sentenses_paraphrased
Dataset Card for "books3_basic_sentenses_paraphrased"
More Information needed
ru-paraphrase-NMT-Leipzig-cleaned
Dataset Description
The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics.
The data structure is saved.
Have been deleted:
Paraphrases that have cosine LABSE similarity with source sentences < 0.75.
Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation)
Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity > 0.6… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.amr-true-paraphrases
True Paraphrases Test Set
The True Paraphrases sentence/phrase pairs derived from the AMR Annotation Guidelines. It was introduced as part of the PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models.
For more details, refer to the original paper that was presented at COLING 2025.
Citation
If you use this dataset, please cite it using the following BibTeX entry:
@inproceedings{michail-etal-2025-paraphrasus,
title = "{PARAPHRASUS}: A… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/amr-true-paraphrases.alpaca_llama3.18b_em_paraphrased_divergence_ratioFOLIO_by_paraphrased_gpt4popqa_full_w_paraphrasessts-h-paraphrase-detection
STS-Hard Test Set
The STS-Hard dataset is a paraphrase detection test set derived from the STSBenchmark dataset. It was introduced as part of the PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models. The test set includes the paraphrase label as well as individual annotation labels from two annotators:
P1: The semanticist.
P2: A student annotator.
For more details, refer to the original paper that was presented at COLING 2025.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/sts-h-paraphrase-detection.bookmia_paraphrase_claudeSO-Python_basics_QA-filtered-2023-T5_paraphrased-tanh_scoresnli-training-paraphrase-augmentation
SNLI Training Paraphrase Augmentation
Purpose
This dataset contains new paraphrases created for training paraphrase augmentation during Phase B of the research project.
It was not used as an NLI evaluation set.
It was not used for paraphrase consistency evaluation.
It is separate from the published SNLI Paraphrase Bank used for evaluation. The generation records report zero collisions with that evaluation bank.
The CSV contains only the new augmentation rows. It… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/snli-training-paraphrase-augmentation.OpenHermes-paraphrased-headlines-2017-2019-eval-setpaul-paraphrased-questionmnli-training-paraphrase-augmentation
MNLI Training Paraphrase Augmentation
Purpose
This dataset contains new paraphrases created for training paraphrase augmentation during Phase B of the research project.
It was not used as an NLI evaluation set.
It was not used for paraphrase consistency evaluation.
It is separate from the published MNLI Paraphrase Bank used for evaluation. The generation records report zero collisions with that evaluation bank.
The CSV contains only the new augmentation rows. It… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/mnli-training-paraphrase-augmentation.FOLIO_by_paraphrased_gpt3.5test_paraphrasecluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid
