datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MNLP_M1_Preference_dpo_dataset
M1 Preference Data for DPO
Dataset Description
This dataset contains processed M1 preference data for DPO training.
Created by: CS-552 Stochastic Parrots Team
Date: May 24, 2025
Version: 1.0
Number of examples: 17615
Dataset Source
This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.INFH-6000Q-dpo-preference-dataset
INFH-6000Q DPO Preference Dataset
This dataset contains the final preference pairs used for the Direct Preference Optimization assignment in this repository.
Source
Base instruction source: GAIR/lima
Candidate generator: local Qwen/Qwen2.5-7B-Instruct
Preference ranker: local llm-blender/PairRM
Construction Pipeline
Sample 50 instructions from the local LIMA training split with seed 42.
Generate 5 candidate responses per instruction with Qwen2.5-7B-Instruct.… See the full description on the dataset page: https://huggingface.co/datasets/ITBill/INFH-6000Q-dpo-preference-dataset.
