datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
universal-preference-hijacking-datasets
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image.
Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images.
This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.SRPO_RL_datasets
SRPO Dataset: Reflection-Aware RL Training Data
This repository provides the multimodal reasoning dataset used in the paper:
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
We release two versions of the dataset:
39K version (modified_39Krelease.jsonl + images.zip)
Enhanced 47K+ version (47K_release_plus.jsonl + 47K_release_plus.zip)
Both follow the same unified format, containing multimodal (image–text) reasoning data with self-reflection… See the full description on the dataset page: https://huggingface.co/datasets/bruce360568/SRPO_RL_datasets.
