datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task045_miscellaneous_sentence_paraphrasing
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task045_miscellaneous_sentence_paraphrasing
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task045_miscellaneous_sentence_paraphrasing.paraphrasing-french
Attribution
MTEB-format derivative of ismailiismail/paraphrasing_french. Query = phrase; corpus = paraphrase.
task1288_glue_mrpc_paraphrasing
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1288_glue_mrpc_paraphrasing
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1288_glue_mrpc_paraphrasing.squad_paraphrasing_5task177_para-nmt_paraphrasing
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task177_para-nmt_paraphrasing
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task177_para-nmt_paraphrasing.paraphrasing-preferences-orpo-dpo
Paraphrasing Preference Dataset
A preference dataset for training paraphrase models via DPO, RLHF, or ORPO. Each example contains a source text, a task-specific prompt, and a chosen/rejected paraphrase pair ranked by a composite quality score.
Dataset Summary
Train
Val
Total
Examples
852
95
947
Sources: Quora questions (571), SQuAD 2.0 sentences (218), CNN News sentences (158). The val split is stratified by excellent_in, category, and binned total_delta.… See the full description on the dataset page: https://huggingface.co/datasets/alecccdd/paraphrasing-preferences-orpo-dpo.Identification-of-paraphrasing
🇰🇿 Identification of Paraphrasing in Kazakh Context
Dataset Summary
Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
2,000
Total Words (approx.)
184,465
Avg. Words per Sample
92
Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.Persian-Text-Paraphrasingparaphrasing_french
Dataset Card for "paraphrasing"
More Information needed
value-systems-in-llms-paraphrasing-and-profile-elicitation
Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness
(Versión en español más abajo.)
Do large language models give stable answers to the same forced-choice question
when the prompt is perturbed in ways that do not change its meaning — and does
assigning them a personality or value profile change those answers?
This dataset contains the full material of that experiment: the 9,350 prompts,
the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.paraphrasing_french_5000
Dataset Card for "paraphrasing_french_5000"
More Information needed
paragraphss_paraphrasing
Dataset Card for "paragraphss_paraphrasing"
More Information needed
multi_paraphrasing_french
Dataset Card for "multi_paraphrasing_french"
More Information needed
flan_combined_task1287_glue_qqp_paraphrasingparaphrasingflan_combined_task1288_glue_mrpc_paraphrasingflan_combined_task177_para-nmt_paraphrasing
