datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ru-paraphrase-NMT-Leipzig
Dataset Card for cointegrated/ru-paraphrase-NMT-Leipzig
Dataset Summary
The dataset contains 1 million Russian sentences and their automatically generated paraphrases.
It was created by David Dale (@cointegrated) by translating the rus-ru_web-public_2019_1M corpus from the Leipzig collection into English and back into Russian. A fraction of the resulting paraphrases are invalid, and should be filtered out.
The blogpost "Перефразирование русских текстов: корпуса, модели… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/ru-paraphrase-NMT-Leipzig.task275_enhanced_wsc_paraphrase_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task275_enhanced_wsc_paraphrase_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task275_enhanced_wsc_paraphrase_generation.ru_paraphraser
Dataset Card for ParaPhraser
Dataset Summary
ParaPhraser is a news headlines corpus annotated according to the following schema:
1: precise paraphrases
0: near paraphrases
-1: non-paraphrases
The Plus part is also available.
It contains clusters of news headline paraphrases labeled automatically by a fine-tuned paraphrase detection BERT model.In order to load it:
from datasets import load_dataset
corpus = load_dataset('merionum/ru_paraphraser', data_files='plus.jsonl')… See the full description on the dataset page: https://huggingface.co/datasets/merionum/ru_paraphraser.task400_paws_paraphrase_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task400_paws_paraphrase_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task400_paws_paraphrase_classification.autoencoder-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models.
It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses).
The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.machine-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools.
It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses).
The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.task442_com_qa_paraphrase_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task442_com_qa_paraphrase_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task442_com_qa_paraphrase_question_generation.style-aware-paraphraser-author-bank-reddit
Style-Aware Paraphraser — Reddit Author Targets Bank
A bank of 12 000 anonymous Reddit authors, each represented by 16 exemplar
comments plus 5 Mistral-7B paraphrases of each. This is what feeds the
target-style side of our paraphraser: pick a row, pass reference_text
and paraphrase_reference_text to
rrivera1849/style-aware-paraphraser-mistral7b,
and the model will rewrite any machine text in that author's style.
Reddit usernames are not included; the bank carries only the… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-author-bank-reddit.style-aware-paraphraser-outputs
Style-Aware Paraphraser Outputs
⚠️ Intended for evaluating machine-text detectors, not for training them.
Each row is a (human, machine, adversarial-paraphrase) triple drawn from a
specific evaluation slice of three domains. The author bank that produced
these outputs overlaps with the rows here, so training a detector against
this data would not generalize. Use it to score a detector, not to fit one.
Final outputs of the style-aware paraphraser from
Attacks on Machine-Text… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-outputs.autoregressive-paraphrase-dataset
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset.nlp-paraphrases-20k-prunedThe dataset contains 20703 records. The dataset was created by removing all dataset items from the original 27k dataset that had a BLEU score 0 or more than 0.3388.
paws-x_fr_prompt_paraphrase_generation
paws-x_fr_prompt_paraphrase_generation
Summary
paws-x_fr_prompt_paraphrase_generation is a subset of the Dataset of French Prompts (DFP).It contains 562,728 rows that can be used for a paraphrase generation task.The original data (without prompts) comes from the dataset paws-x by Yang et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/paws-x_fr_prompt_paraphrase_generation.bangla-health-related-paraphrased-dataset
Dataset Card for "BanglaHealthParaphrase"
BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.gerlayqa-bgb-paraphrased
GerLayQA-BGB Paraphrased 🇩🇪⚖️
Dataset Description
This is a paraphrased and restructured version of the GerLayQA BGB (Bürgerliches Gesetzbuch / German Civil Code) dataset, specifically prepared for fine-tuning large language models on German civil law question-answering tasks.
Key Features
5,255 high-quality QA pairs about German Civil Law (BGB)
Paraphrased questions to remove plagiarism while maintaining legal accuracy
Structured 7-section answers following… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-bgb-paraphrased.gerlayqa-combined-paraphrased
GerLayQA Combined Paraphrased 🇩🇪⚖️
Dataset Description
This is a combined, shuffled dataset merging both the BGB (civil law) and StGB (criminal law) paraphrased German legal QA datasets. All examples are paraphrased and restructured by GPT-5 for fine-tuning large language models on German legal question-answering tasks.
Key Features
6,462 high-quality QA pairs covering both German Civil and Criminal Law
Combined coverage: BGB (Bürgerliches Gesetzbuch) + StGB… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-combined-paraphrased.sentence-paraphrases
Sentence Paraphrases Dataset
This dataset is a curated collection of sentence-length paraphrases derived from two primary sources:
humarin/chatgpt-paraphrases
xwjzds/paraphrase_collections.
Dataset Details
Dataset Description
The dataset is structured to provide pairs of sentences from an original text and its paraphrase(s). For each entry:
The "text" field contains the least readable paraphrase.
The "paraphrase" field contains the most readable paraphrase.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/sentence-paraphrases.gerlayqa-stgb-paraphrased
GerLayQA-StGB Paraphrased 🇩🇪⚖️
Dataset Description
This is a paraphrased and restructured version of the GerLayQA StGB (Strafgesetzbuch / German Criminal Code) dataset, specifically prepared for fine-tuning large language models on German criminal law question-answering tasks.
Key Features
1,207 high-quality QA pairs about German Criminal Law (StGB)
Paraphrased questions to remove plagiarism while maintaining legal accuracy
Structured 7-section answers… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-stgb-paraphrased.gsm8k-paraphrase-deepseek
GSM8K paraphrase variants (DeepSeek-v4-flash)
Paraphrase ablations over the GSM8K-test split, formatted to match the
openai/gsm8k schema: a question string and an answer string ending with
#### N.
The first three configs cover the same 660-item subsample used to
contaminate the m2_mixed model
(random.Random(42).sample(range(1319), 660)). The fourth config,
pq_complement, covers the remaining 659 held-out items (the complement of
that subsample), so pq ∪ pq_complement… See the full description on the dataset page: https://huggingface.co/datasets/onnookk/gsm8k-paraphrase-deepseek.ru-paraphrase-NMT-Leipzig-cleaned
Dataset Description
The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics.
The data structure is saved.
Have been deleted:
Paraphrases that have cosine LABSE similarity with source sentences < 0.75.
Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation)
Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity > 0.6… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.aya-telugu-paraphrase
Summary
aya-telugu-paraphrase is an open source dataset of instruct-style records generated from the Telugu split of ai4bharat/IndicXParaphrase dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-paraphrase.DPO-arxiv_paraphraseGPT-Paraphrases
GPT-Paraphrases dataset
This dataset contains text passages and their paraphrases generated using the GPT-3 language model. The paraphrases are designed to be semantically equivalent to the original text, but with different wording and structure.
The dataset includes text formatted in JSON and is in English.
Dataset Statistics
Number of text passages: Not specified in the information you provided.
Source of text passages: Not specified in the information you provided.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/GPT-Paraphrases.wikipedia-paragraphs-direct-paraphrases
Wikipedia Paragraphs Direct Paraphrases
Paraphrases of Wikipedia paragraphs using AI large language models.
Paragraphs from agentlans/wikipedia-paragraphs-complete sample_k10000 and sample_k50000 splits
Paraphrased using Qwen/Qwen3.5-9B and a distilled Qwen/Qwen3-4B-Instruct-2507 with the following prompt:
Rewrite the following paragraph entirely in your own words while preserving every fact, detail, meaning, nuance, and level of specificity. Do not add, remove, reinterpret… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-direct-paraphrases.aya-processed-central-kurdish-paraphrase
Aya Processed Central Kurdish Paraphrase Identification Dataset
Dataset Description
This dataset contains 49,401 preprocessed Central Kurdish paraphrase identification samples in conversational format, derived from the Aya Collection Language Split.
Languages
Central Kurdish (ckb)
Dataset Structure
The dataset contains the following columns:
id: Unique identifier for each sample
prompt: Conversational format as JSON array containing the… See the full description on the dataset page: https://huggingface.co/datasets/shiima/aya-processed-central-kurdish-paraphrase.parabank-paraphrase-labeled
ParaBank Paraphrase Labeled (Style-Conditioned)
200,001-row style-labeled paraphrase dataset derived from ParaBank v2, balanced at
66,667 rows per style: [FORMAL], [CASUAL] (informal), [POLITE].
Columns
reference: source English sentence
paraphrase_1: best-ranked paraphrase from ParaBank v2
style_label: 0=polite, 1=formal, 2=informal
plus raw formality/politeness classifier labels and confidence scores
Data Attribution
Derived from ParaBank 2:
J.… See the full description on the dataset page: https://huggingface.co/datasets/shaheeeeeeeeem/parabank-paraphrase-labeled.vietnamese-author-styles-paraphrasedPolish-paraphrases-12K-synthetic12K sentence paraphrases in Polish.
First, lot of various sentences were generated using ChatGPT 5.2 and Gemini 3 Thinking, then all of them were paraphrased by GPT-4o.
Sentences range from short to medium to long, and are not only random general sentences but also cover tech topics and topics related to universities and student life.
Please mind the license terms for content generated by the AI models mentioned when using this dataset. ChatGPT itself belongs to OpenAI, Gemini to Google.
12… See the full description on the dataset page: https://huggingface.co/datasets/Wojtekb30/Polish-paraphrases-12K-synthetic.
