CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cointegrated /ru-paraphrase-NMT-Leipzig Dataset Card for cointegrated/ru-paraphrase-NMT-Leipzig Dataset Summary The dataset contains 1 million Russian sentences and their automatically generated paraphrases. It was created by David Dale (@cointegrated) by translating the rus-ru_web-public_2019_1M corpus from the Leipzig collection into English and back into Russian. A fraction of the resulting paraphrases are invalid, and should be filtered out. The blogpost "Перефразирование русских текстов: корпуса, модели… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/ru-paraphrase-NMT-Leipzig.text-generation100K<n<1M11 likes524 downloads4y agoHugging Face02Lots-of-LoRAs /task275_enhanced_wsc_paraphrase_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task275_enhanced_wsc_paraphrase_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task275_enhanced_wsc_paraphrase_generation.texttext-generation1K<n<10K2 likes368 downloads2y agoHugging Face03merionum /ru_paraphraser Dataset Card for ParaPhraser Dataset Summary ParaPhraser is a news headlines corpus annotated according to the following schema: 1: precise paraphrases 0: near paraphrases -1: non-paraphrases The Plus part is also available. It contains clusters of news headline paraphrases labeled automatically by a fine-tuned paraphrase detection BERT model.In order to load it: from datasets import load_dataset corpus = load_dataset('merionum/ru_paraphraser', data_files='plus.jsonl')… See the full description on the dataset page: https://huggingface.co/datasets/merionum/ru_paraphraser.texttext-classification1K<n<10K13 likes303 downloads4y agoHugging Face04Lots-of-LoRAs /task400_paws_paraphrase_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task400_paws_paraphrase_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task400_paws_paraphrase_classification.texttext-generation1K<n<10K0 likes231 downloads2y agoHugging Face05jpwahle /autoencoder-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models. It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses). The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.tabulartext-classification1M<n<10M2 likes141 downloads1y agoHugging Face06jpwahle /machine-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools. It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses). The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.texttext-classification100K<n<1M7 likes131 downloads1y agoHugging Face07Lots-of-LoRAs /task442_com_qa_paraphrase_question_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task442_com_qa_paraphrase_question_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task442_com_qa_paraphrase_question_generation.texttext-generation1K<n<10K0 likes118 downloads2y agoHugging Face08rrivera1849 /style-aware-paraphraser-author-bank-reddit Style-Aware Paraphraser — Reddit Author Targets Bank A bank of 12 000 anonymous Reddit authors, each represented by 16 exemplar comments plus 5 Mistral-7B paraphrases of each. This is what feeds the target-style side of our paraphraser: pick a row, pass reference_text and paraphrase_reference_text to rrivera1849/style-aware-paraphraser-mistral7b, and the model will rewrite any machine text in that author's style. Reddit usernames are not included; the bank carries only the… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-author-bank-reddit.texttext-generation10K<n<100K1 likes55 downloads3mo agoHugging Face09rrivera1849 /style-aware-paraphraser-outputs Style-Aware Paraphraser Outputs ⚠️ Intended for evaluating machine-text detectors, not for training them. Each row is a (human, machine, adversarial-paraphrase) triple drawn from a specific evaluation slice of three domains. The author bank that produced these outputs overlaps with the rows here, so training a detector against this data would not generalize. Use it to score a detector, not to fit one. Final outputs of the style-aware paraphraser from Attacks on Machine-Text… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-outputs.texttext-classification10K<n<100K0 likes50 downloads3mo agoHugging Face10jpwahle /autoregressive-paraphrase-dataset Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset.texttext-classification100K<n<1M1 likes48 downloads4y agoHugging Face11skupina-7 /nlp-paraphrases-20k-prunedThe dataset contains 20703 records. The dataset was created by removing all dataset items from the original 27k dataset that had a BLEU score 0 or more than 0.3388. texttext-generation10K<n<100K0 likes40 downloads3y agoHugging Face12CATIE-AQ /paws-x_fr_prompt_paraphrase_generation paws-x_fr_prompt_paraphrase_generation Summary paws-x_fr_prompt_paraphrase_generation is a subset of the Dataset of French Prompts (DFP).It contains 562,728 rows that can be used for a paraphrase generation task.The original data (without prompts) comes from the dataset paws-x by Yang et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/paws-x_fr_prompt_paraphrase_generation.texttext-generation100K<n<1M1 likes38 downloads1y agoHugging Face13faisal4590aziz /bangla-health-related-paraphrased-dataset Dataset Card for "BanglaHealthParaphrase" BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.tabulartext-generation100K<n<1M2 likes33 downloads1y agoHugging Face14DomainLLM /gerlayqa-bgb-paraphrased GerLayQA-BGB Paraphrased 🇩🇪⚖️ Dataset Description This is a paraphrased and restructured version of the GerLayQA BGB (Bürgerliches Gesetzbuch / German Civil Code) dataset, specifically prepared for fine-tuning large language models on German civil law question-answering tasks. Key Features 5,255 high-quality QA pairs about German Civil Law (BGB) Paraphrased questions to remove plagiarism while maintaining legal accuracy Structured 7-section answers following… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-bgb-paraphrased.textquestion-answering1K<n<10K0 likes32 downloads1y agoHugging Face15DomainLLM /gerlayqa-combined-paraphrased GerLayQA Combined Paraphrased 🇩🇪⚖️ Dataset Description This is a combined, shuffled dataset merging both the BGB (civil law) and StGB (criminal law) paraphrased German legal QA datasets. All examples are paraphrased and restructured by GPT-5 for fine-tuning large language models on German legal question-answering tasks. Key Features 6,462 high-quality QA pairs covering both German Civil and Criminal Law Combined coverage: BGB (Bürgerliches Gesetzbuch) + StGB… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-combined-paraphrased.textquestion-answering1K<n<10K0 likes27 downloads1y agoHugging Face16agentlans /sentence-paraphrases Sentence Paraphrases Dataset This dataset is a curated collection of sentence-length paraphrases derived from two primary sources: humarin/chatgpt-paraphrases xwjzds/paraphrase_collections. Dataset Details Dataset Description The dataset is structured to provide pairs of sentences from an original text and its paraphrase(s). For each entry: The "text" field contains the least readable paraphrase. The "paraphrase" field contains the most readable paraphrase.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/sentence-paraphrases.texttext-generation100K<n<1M3 likes24 downloads2y agoHugging Face17DomainLLM /gerlayqa-stgb-paraphrased GerLayQA-StGB Paraphrased 🇩🇪⚖️ Dataset Description This is a paraphrased and restructured version of the GerLayQA StGB (Strafgesetzbuch / German Criminal Code) dataset, specifically prepared for fine-tuning large language models on German criminal law question-answering tasks. Key Features 1,207 high-quality QA pairs about German Criminal Law (StGB) Paraphrased questions to remove plagiarism while maintaining legal accuracy Structured 7-section answers… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-stgb-paraphrased.textquestion-answering1K<n<10K0 likes24 downloads1y agoHugging Face18onnookk /gsm8k-paraphrase-deepseek GSM8K paraphrase variants (DeepSeek-v4-flash) Paraphrase ablations over the GSM8K-test split, formatted to match the openai/gsm8k schema: a question string and an answer string ending with #### N. The first three configs cover the same 660-item subsample used to contaminate the m2_mixed model (random.Random(42).sample(range(1319), 660)). The fourth config, pq_complement, covers the remaining 659 held-out items (the complement of that subsample), so pq ∪ pq_complement… See the full description on the dataset page: https://huggingface.co/datasets/onnookk/gsm8k-paraphrase-deepseek.texttext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face19fyaronskiy /ru-paraphrase-NMT-Leipzig-cleaned Dataset Description The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics. The data structure is saved. Have been deleted: Paraphrases that have cosine LABSE similarity with source sentences < 0.75. Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation) Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity > 0.6… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.tabulartext-generation100K<n<1M2 likes23 downloads1y agoHugging Face20SuryaKrishna02 /aya-telugu-paraphrase Summary aya-telugu-paraphrase is an open source dataset of instruct-style records generated from the Telugu split of ai4bharat/IndicXParaphrase dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-paraphrase.texttext-generation1K<n<10K4 likes22 downloads3y agoHugging Face21Mxode /DPO-arxiv_paraphrasetexttext-generation100K<n<1M2 likes21 downloads1y agoHugging Face22prithivMLmods /GPT-Paraphrases GPT-Paraphrases dataset This dataset contains text passages and their paraphrases generated using the GPT-3 language model. The paraphrases are designed to be semantically equivalent to the original text, but with different wording and structure. The dataset includes text formatted in JSON and is in English. Dataset Statistics Number of text passages: Not specified in the information you provided. Source of text passages: Not specified in the information you provided.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/GPT-Paraphrases.texttext-generation100K<n<1M1 likes20 downloads2y agoHugging Face23agentlans /wikipedia-paragraphs-direct-paraphrases Wikipedia Paragraphs Direct Paraphrases Paraphrases of Wikipedia paragraphs using AI large language models. Paragraphs from agentlans/wikipedia-paragraphs-complete sample_k10000 and sample_k50000 splits Paraphrased using Qwen/Qwen3.5-9B and a distilled Qwen/Qwen3-4B-Instruct-2507 with the following prompt: Rewrite the following paragraph entirely in your own words while preserving every fact, detail, meaning, nuance, and level of specificity. Do not add, remove, reinterpret… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-direct-paraphrases.texttext-generation10K<n<100K0 likes18 downloads3mo agoHugging Face24shiima /aya-processed-central-kurdish-paraphrase Aya Processed Central Kurdish Paraphrase Identification Dataset Dataset Description This dataset contains 49,401 preprocessed Central Kurdish paraphrase identification samples in conversational format, derived from the Aya Collection Language Split. Languages Central Kurdish (ckb) Dataset Structure The dataset contains the following columns: id: Unique identifier for each sample prompt: Conversational format as JSON array containing the… See the full description on the dataset page: https://huggingface.co/datasets/shiima/aya-processed-central-kurdish-paraphrase.texttext-generation10K<n<100K0 likes17 downloads8mo agoHugging Face25shaheeeeeeeeem /parabank-paraphrase-labeled ParaBank Paraphrase Labeled (Style-Conditioned) 200,001-row style-labeled paraphrase dataset derived from ParaBank v2, balanced at 66,667 rows per style: [FORMAL], [CASUAL] (informal), [POLITE]. Columns reference: source English sentence paraphrase_1: best-ranked paraphrase from ParaBank v2 style_label: 0=polite, 1=formal, 2=informal plus raw formality/politeness classifier labels and confidence scores Data Attribution Derived from ParaBank 2: J.… See the full description on the dataset page: https://huggingface.co/datasets/shaheeeeeeeeem/parabank-paraphrase-labeled.tabulartext-generation100K<n<1M0 likes14 downloads3mo agoHugging Face26DiligentPenguinn /vietnamese-author-styles-paraphrasedtexttext-generationn<1K0 likes9 downloads1y agoHugging Face27Wojtekb30 /Polish-paraphrases-12K-synthetic12K sentence paraphrases in Polish. First, lot of various sentences were generated using ChatGPT 5.2 and Gemini 3 Thinking, then all of them were paraphrased by GPT-4o. Sentences range from short to medium to long, and are not only random general sentences but also cover tech topics and topics related to universities and student life. Please mind the license terms for content generated by the AI models mentioned when using this dataset. ChatGPT itself belongs to OpenAI, Gemini to Google. 12… See the full description on the dataset page: https://huggingface.co/datasets/Wojtekb30/Polish-paraphrases-12K-synthetic.tabulartext-generation1K<n<10K0 likes8 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.