datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paraphrasing-french
Attribution
MTEB-format derivative of ismailiismail/paraphrasing_french. Query = phrase; corpus = paraphrase.
paraphrasing-preferences-orpo-dpo
Paraphrasing Preference Dataset
A preference dataset for training paraphrase models via DPO, RLHF, or ORPO. Each example contains a source text, a task-specific prompt, and a chosen/rejected paraphrase pair ranked by a composite quality score.
Dataset Summary
Train
Val
Total
Examples
852
95
947
Sources: Quora questions (571), SQuAD 2.0 sentences (218), CNN News sentences (158). The val split is stratified by excellent_in, category, and binned total_delta.… See the full description on the dataset page: https://huggingface.co/datasets/alecccdd/paraphrasing-preferences-orpo-dpo.value-systems-in-llms-paraphrasing-and-profile-elicitation
Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness
(Versión en español más abajo.)
Do large language models give stable answers to the same forced-choice question
when the prompt is perturbed in ways that do not change its meaning — and does
assigning them a personality or value profile change those answers?
This dataset contains the full material of that experiment: the 9,350 prompts,
the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.paraphrasing
