datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
universal-preference-hijacking-datasets
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image.
Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images.
This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.gsm8k_preference_dataset_it_1
GSM8K Iteration 1
Overview
This dataset is derived from the GSM8K training set questions. The process to create this dataset involved the following steps:
Initial Prompting: Each question from the GSM8K train set was initially answered by the Mistral model.
Filtering Incorrect Answers: Incorrect responses were filtered out.
Refinement: The model was prompted to refine its answers based on the incorrect responses.
Final Filtering: The refined responses were filtered again… See the full description on the dataset page: https://huggingface.co/datasets/August4293/gsm8k_preference_dataset_it_1.Self_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.
