datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
universal-preference-hijacking-datasets
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image.
Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images.
This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.m1_preference_data_cleaned
EPFL M1 MCQ Dataset (Cleaned)
This dataset contains 645 multiple-choice questions extracted and cleaned from EPFL M1 preference data. Each question has exactly 4 options (A, B, C, D) with balanced sampling when original questions had more options.
Dataset Statistics
Total Questions: 645
Format: Multiple choice questions with exactly 4 options
Domain: Computer Science and Engineering
Source: EPFL M1 preference data
Answer Distribution: A: 204, B: 145, C: 147, D: 149… See the full description on the dataset page: https://huggingface.co/datasets/albertfares/m1_preference_data_cleaned.bias_reduce_preference_data
Dataset Card for Persona-Aware Preference Dataset
Dataset Description
This is a Direct Preference Optimization (DPO) dataset designed to train language models to produce high-quality, context-aware responses when given user demographic information (persona). Each example pairs a user prompt prefixed with a demographic persona description with a chosen (preferred) response and a rejected (dispreferred) response.
The dataset is intended to support alignment research focused… See the full description on the dataset page: https://huggingface.co/datasets/groupfairnessllm/bias_reduce_preference_data.gsm8k_preference_dataset_it_1
GSM8K Iteration 1
Overview
This dataset is derived from the GSM8K training set questions. The process to create this dataset involved the following steps:
Initial Prompting: Each question from the GSM8K train set was initially answered by the Mistral model.
Filtering Incorrect Answers: Incorrect responses were filtered out.
Refinement: The model was prompted to refine its answers based on the incorrect responses.
Final Filtering: The refined responses were filtered again… See the full description on the dataset page: https://huggingface.co/datasets/August4293/gsm8k_preference_dataset_it_1.bitext_wealth_management_preference_dataThis is a dataset created from the train split of the bitext-wealth_management-llm-chatbot dataset. The chosen response is the ground truth response. The rejected response is the ony selected by gpt-4o out of a list of candidate responses from an sft trained model.
Self_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.
