datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jitteredwebsites-merged-224-paraphrasedtask275_enhanced_wsc_paraphrase_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task275_enhanced_wsc_paraphrase_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task275_enhanced_wsc_paraphrase_generation.fineweb-edu-paraphrasedru_paraphraser
Dataset Card for ParaPhraser
Dataset Summary
ParaPhraser is a news headlines corpus annotated according to the following schema:
1: precise paraphrases
0: near paraphrases
-1: non-paraphrases
The Plus part is also available.
It contains clusters of news headline paraphrases labeled automatically by a fine-tuned paraphrase detection BERT model.In order to load it:
from datasets import load_dataset
corpus = load_dataset('merionum/ru_paraphraser', data_files='plus.jsonl')… See the full description on the dataset page: https://huggingface.co/datasets/merionum/ru_paraphraser.ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.enwiki20250301_paraphrase_multilingual_minilm_l12_v2task400_paws_paraphrase_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task400_paws_paraphrase_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task400_paws_paraphrase_classification.chatgpt-paraphrasesThis is a dataset of paraphrases created by ChatGPT.
Model based on this dataset is avaible: model
We used this prompt to generate paraphrases
Generate 5 similar paraphrases for this question, show it like a numbered list without commentaries: {text}
This dataset is based on the Quora paraphrase question, texts from the SQUAD 2.0 and the CNN news dataset.
We generated 5 paraphrases for each sample, totally this dataset has about 420k data rows. You can make 30 rows from a row from… See the full description on the dataset page: https://huggingface.co/datasets/humarin/chatgpt-paraphrases.paraphrase-qa日本語Wikipedia中のテキストを元に言い換えを生成し、その言い換えを元にクエリと回答をLLMに生成させたデータセットです。
出力にライセンス的な制約があるモデルを利用していないことと、元データとして日本語Wikipediaを利用していることから、CC-BY-SA 4.0ライセンスのもとでの配布とします。
autoencoder-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models.
It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses).
The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.paraphrases_mrpcmodular-sentence-encoders-paraphrase
Multi-parallel paraphrase corpus (23 languages)
The contrastive training data of Modular Sentence Encoders: Separating Language
Specialization from Cross-Lingual Alignment
(ACL 2025). Five English paraphrase datasets, each translated into 22 further
languages, published so that the paper's sentence-encoder and alignment stages
can be reproduced without re-running the translation.
Most of this corpus is machine-translated. Only the English columns are
original human-written text;… See the full description on the dataset page: https://huggingface.co/datasets/yoh/modular-sentence-encoders-paraphrase.machine-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools.
It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses).
The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.paraphrases_pawstask442_com_qa_paraphrase_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task442_com_qa_paraphrase_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task442_com_qa_paraphrase_question_generation.paraphrasejitteredwebsites-merged-224-paraphrased-pairedtoxigen-paraphrased
toxigen-paraphrased
Paraphrased version of the ToxiGen (annotated) dataset. Each text has been paraphrased while preserving its original toxicity label.
Model
Paraphrases were generated using Qwen/Qwen3-30B-A3B-Instruct-2507.
Original Dataset
Source: skg/toxigen-data (annotated)
Task: Toxicity Detection
Classes: 2 (0 = non-toxic, 1 = toxic)
Dataset Structure
Split
Examples
train
7,168
validation
1,792
test
940
Columns… See the full description on the dataset page: https://huggingface.co/datasets/concretejungles/toxigen-paraphrased.llm-paraphrases
LLM Generated Paraphrases Dataset
A large-scale synthetically generated paraphrase dataset containing sentence pairs with balanced positive and negative examples across varied domains and writing styles.
Dataset Details
Dataset Description
Name: llm-paraphrases
Summary: A synthetic paraphrase dataset generated using large language models, designed for training embedding models for semantic caching and paraphrase detection. Each example contains a pair of… See the full description on the dataset page: https://huggingface.co/datasets/redis/llm-paraphrases.sbert-paraphrase-data
BEE-spoke-data/sbert-paraphrase-data
Paraphrase data from sentence-transformers
contents
default
No.
Filename
1
yahoo_answers_title_question.jsonl
2
squad_pairs.jsonl
3
eli5_question_answer.jsonl
4
WikiAnswers_pairs.jsonl
5
stackexchange_duplicate_questions_title_title.jsonl
6
TriviaQA_pairs.jsonl
7
stackexchange_duplicate_questions.jsonl
8
sentence-compression.jsonl
9
AllNLI_2cols.jsonl
10
NQ-train_pairs.jsonl
11… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sbert-paraphrase-data.ger-backtrans-paraphrase
German Backtranslated Paraphrase Dataset
This is a dataset of more than 21 million German paraphrases.
These are text pairs that have the same meaning but are expressed with different words.
The source of the paraphrases are different parallel German / English text corpora.
The English texts were machine translated back into German to obtain the paraphrases.
This dataset can be used for example to train semantic text embeddings.
To do this, for example, SentenceTransformers
and the… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/ger-backtrans-paraphrase.jailbreakbench-paraphrase-2025-08
JailbreakBench Paraphrase Dataset (2025-08)
This dataset contains 115 paraphrased prompts from the JailbreakBench dataset, created as part of research on semantic entropy-based jailbreak detection and its robustness to paraphrasing.
Dataset Description
Overview
This dataset was created to test the robustness of jailbreak detection methods (particularly semantic entropy) to paraphrased inputs. The paraphrases were generated from the original JailbreakBench… See the full description on the dataset page: https://huggingface.co/datasets/DhruvTre/jailbreakbench-paraphrase-2025-08.ltgen-wiki-paraphrased-Humanized-19999tr-paraphrase-opensubtitles2018WikiMIA_paraphrased_perturbed
📘 WikiMIA paraphrased and perturbed versions
The WikiMIA dataset serves as a benchmark designed to evaluate membership inference attack (MIA) methods, specifically in detecting pretraining data from extensive large language models.
It is originally constructed by Shi et al. (see the original data repo for more details).
The authors studied a paraphrased setting in their paper, where instead of detecting verbatim training texts, the goal is to detect (slightly) paraphrased… See the full description on the dataset page: https://huggingface.co/datasets/zjysteven/WikiMIA_paraphrased_perturbed.sst5-paraphrased
sst5-paraphrased
Paraphrased version of the Stanford Sentiment Treebank v5 (SST-5) dataset. Each sentence has been paraphrased while preserving its original fine-grained sentiment label.
Model
Paraphrases were generated using Qwen/Qwen3-30B-A3B-Instruct-2507.
Original Dataset
Source: SetFit/sst5
Task: Fine-grained Sentiment Classification
Classes: 5 (0 = very negative, 1 = negative, 2 = neutral, 3 = positive, 4 = very positive)
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/concretejungles/sst5-paraphrased.chatgpt-paraphrases-simpleThis dataset is simplified version of ChatGPT Paraphrases. And aims to take away the pain of expanding original dataset into unique paraphrase pairs.
Structure:
Dataset is not divided into train/test split. And contains 6.3 million unique paraphrases(6x5x420000/2 = 6.3 million). Dataset contains following 2 columns-
s1 - Sentence
s2 - Paraphrase
Original Dataset Structure:
The original dataset has following 4 columns-
text - 420k Unique sentence
paraphrases - List of 5 unique… See the full description on the dataset page: https://huggingface.co/datasets/sharad/chatgpt-paraphrases-simple.LLM-Oasis_paraphrase_generation
Babelscape/LLM-Oasis_paraphrase_generation
Dataset Description
LLM-Oasis_paraphrase_generation is part of the LLM-Oasis suite and contains paraphrases generated from a set of claims extracted from a Wikipedia passage.
This dataset supports the paraphrase generation step described in Section 3.3 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis.
Features
title: The title… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_paraphrase_generation.style-aware-paraphraser-author-bank-reddit
Style-Aware Paraphraser — Reddit Author Targets Bank
A bank of 12 000 anonymous Reddit authors, each represented by 16 exemplar
comments plus 5 Mistral-7B paraphrases of each. This is what feeds the
target-style side of our paraphraser: pick a row, pass reference_text
and paraphrase_reference_text to
rrivera1849/style-aware-paraphraser-mistral7b,
and the model will rewrite any machine text in that author's style.
Reddit usernames are not included; the bank carries only the… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-author-bank-reddit.farsi_paraphrase_detection
Dataset Card for "farsi_paraphrase_detection"
More Information needed
